feat(durable-jobs): support successful rescheduling - #10716
Conversation
Add explicit running and reschedule-requested outcomes, route successful reschedules through reset-capable built-in shards, and preserve retry semantics for failures and unknown statuses. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Pull request overview
Adds a first-class “successful reschedule” terminal outcome to the Durable Jobs execution model, allowing a job execution to complete successfully while requesting a new durable run (with attempt count reset), plus associated executor routing, logging, metrics, and tests across in-memory and journaled shards.
Changes:
- Introduce
DurableJobRunResult.RescheduleAt(...)/DurableJobRunStatus.RescheduleRequestedand route that outcome through an internal reset-rescheduling shard capability. - Rename the “poll again” status/result from
PollAfter/IsPendingtoRunning/IsRunning, and add explicit handling for unknown status values via the failure retry policy. - Add new rescheduling metrics/logging and expand test coverage for rescheduling semantics (including shard reassignment persistence).
Show a summary per file
| File | Description |
|---|---|
| test/Orleans.DurableJobs.Tests/DurableJobs/JobShardManagerTestsRunner.cs | Adds shard reassignment test ensuring successful reschedule persists with reset dequeue count and new run id. |
| test/Orleans.Core.Tests/DurableJobs/ShardExecutorTests.cs | Adds executor tests for rescheduling behavior, legacy shard behavior, and unknown-status retry-policy flow; updates “polling” tests to “running”. |
| test/Orleans.Core.Tests/DurableJobs/JobShardTests.cs | Adds unit test verifying reschedule resets dequeue count and yields a new run id vs retry semantics. |
| test/Orleans.Core.Tests/DurableJobs/DurableJobsInstrumentsTests.cs | Extends metrics assertions to cover reschedule and attempt-duration tagging. |
| test/Orleans.Core.Tests/DurableJobs/DurableJobRunResultTests.cs | New tests for enum numeric stability, serializer member ids, and new result invariants. |
| test/Orleans.Core.Tests/DurableJobs/DurableJobReceiverExtensionTests.cs | Updates tests for renamed “running” semantics and removes redundant null-exception test now covered elsewhere. |
| src/Orleans.DurableJobs/ShardExecutor.Log.cs | Adds a dedicated rescheduling log message. |
| src/Orleans.DurableJobs/ShardExecutor.cs | Implements reschedule handling as a successful terminal outcome (no failure retry policy), plus explicit unknown-status handling. |
| src/Orleans.DurableJobs/JournaledJobShard.cs | Adds reset-rescheduling support by allowing dequeue-count reset during retry/reschedule persistence. |
| src/Orleans.DurableJobs/JobShard.cs | Implements reset-rescheduling for the base in-memory shard and introduces IResettableJobShard. |
| src/Orleans.DurableJobs/IDurableJobReceiverExtension.cs | Updates the receiver extension to return “running” instead of “poll-after” for ongoing executions. |
| src/Orleans.DurableJobs/DurableJobsInstruments.cs | Adds rescheduled counters and tagged attempt-duration recording for reschedules. |
| src/Orleans.DurableJobs/DurableJobRunResult.cs | Adds reschedule result/status, renames polling result/status to “running”, and extends serialized payload with reschedule time. |
| src/api/Orleans.DurableJobs/Orleans.DurableJobs.cs | Updates the public API baseline for the new/renamed result/status members. |
Review details
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
- Files reviewed: 14/14 changed files
- Comments generated: 1
- Review effort level: Lite
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Review details
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
src/Orleans.DurableJobs/ShardExecutor.cs:280
- When a job returns RescheduleRequested on a shard that does not implement IResettableJobShard, the code throws ResettableJobShardNotSupportedException but does not emit any log entry (it is excluded from the inner catch which logs, and the outer catch only enqueues the exception). This can make the failure hard to diagnose in production logs.
case DurableJobRunStatus.RescheduleRequested when result.RescheduleTime is { } rescheduleTime:
if (shard is not IResettableJobShard resettableShard)
{
throw new ResettableJobShardNotSupportedException(shard.GetType());
}
- Files reviewed: 18/18 changed files
- Comments generated: 0 new
- Review effort level: Lite
Use InProgress for active executions which retain their concurrency slot while awaiting the next poll. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2fd6d15a-9067-4a76-9b1f-05c9a38d09cf
There was a problem hiding this comment.
Review details
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
src/Orleans.DurableJobs/JobShard.cs:207
RescheduleJobAsynccreates a newJobRunContextwithretryCount: 0(representing a fresh run), but it reuses the prior run’sRunId. SinceRunIdis documented as “not preserved across retries”, generating a newRunIdhere better matches the reset-rescheduling semantics and avoids leaking the prior run identity into persistence implementations.
var executionGeneration = checked(jobContext.Job.ExecutionGeneration + 1);
var resetContext = new JobRunContext(
jobContext.Job.WithExecutionGeneration(executionGeneration),
jobContext.RunId,
retryCount: 0);
- Files reviewed: 18/18 changed files
- Comments generated: 1
- Review effort level: Lite
Remove the unreleased PollAfter and IsPending aliases so InProgress is the sole active-execution API. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2fd6d15a-9067-4a76-9b1f-05c9a38d09cf
There was a problem hiding this comment.
Review details
Suppressed comments (3)
src/Orleans.DurableJobs/DurableJobRunResult.cs:22
- This change removes the previously existing public
DurableJobRunResult.IsPendingAPI. That’s a source-breaking change for consumers; please keep it as an[Obsolete]forwarding alias toIsInProgressto preserve compatibility while introducing the new naming.
[MemberNotNullWhen(true, nameof(PollAfterDelay))]
public bool IsInProgress => Status == DurableJobRunStatus.InProgress;
src/Orleans.DurableJobs/DurableJobRunResult.cs:87
- The public
DurableJobRunResult.PollAfter(TimeSpan)factory method was removed. This is a source-breaking change; keep it as an[Obsolete]forwarding alias toInProgress(...)so existing callers continue to compile.
public static DurableJobRunResult InProgress(TimeSpan delay)
{
ArgumentOutOfRangeException.ThrowIfLessThanOrEqual(delay, TimeSpan.Zero, nameof(delay));
return new(DurableJobRunStatus.InProgress, delay, null, null);
}
src/Orleans.DurableJobs/DurableJobRunResult.cs:134
DurableJobRunStatus.PollAfter(serialized value 1) appears to have been renamed toInProgresswithout keeping an alias. To avoid breaking existing consumers, keepPollAfteras an[Obsolete]enum member aliasing the same value.
/// The job execution remains in progress and should be polled again after the specified delay.
/// </summary>
InProgress = 1,
- Files reviewed: 18/18 changed files
- Comments generated: 1
- Review effort level: Lite
Expose successful reset-rescheduling directly on the mutable pre-release IJobShard API and remove the auxiliary capability interface. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2fd6d15a-9067-4a76-9b1f-05c9a38d09cf
There was a problem hiding this comment.
Review details
Suppressed comments (5)
src/Orleans.DurableJobs/DurableJobRunResult.cs:87
- The previous
DurableJobRunResult.PollAfter(TimeSpan)factory is no longer present, which is a public API breaking change. Add it back as an[Obsolete]forwarding alias toInProgress(TimeSpan)so existing code continues to compile.
public static DurableJobRunResult InProgress(TimeSpan delay)
{
ArgumentOutOfRangeException.ThrowIfLessThanOrEqual(delay, TimeSpan.Zero, nameof(delay));
return new(DurableJobRunStatus.InProgress, delay, null, null);
}
src/Orleans.DurableJobs/DurableJobRunResult.cs:23
- The previous
DurableJobRunResult.IsPendingmember was removed/renamed toIsInProgress. That’s a source-breaking change for existing callers and contradicts the prior resolution to keep legacy members as obsolete forwarding aliases. ReintroduceIsPendingas an[Obsolete]alias which forwards toIsInProgress.
/// <summary>
/// Gets a value indicating whether the job execution remains in progress and should be polled after a delay.
/// </summary>
[MemberNotNullWhen(true, nameof(PollAfterDelay))]
public bool IsInProgress => Status == DurableJobRunStatus.InProgress;
src/Orleans.DurableJobs/DurableJobRunResult.cs:134
DurableJobRunStatus.PollAfterwas renamed toInProgresswithout keeping a legacy alias. Since the enum is public, removing the old name is source-breaking. KeepPollAfteras an[Obsolete]alias with the same underlying value for compatibility.
/// <summary>
/// The job execution remains in progress and should be polled again after the specified delay.
/// </summary>
InProgress = 1,
src/api/Orleans.DurableJobs/Orleans.DurableJobs.cs:57
- The generated API surface file reflects the removal of
IsPending/PollAfter(...). Once the legacy members are restored as[Obsolete]aliases in the implementation, update/regenerate the API surface so it includes them too (otherwise the published surface remains breaking).
[System.Diagnostics.CodeAnalysis.MemberNotNullWhen(true, "PollAfterDelay")]
public bool IsInProgress { get { throw null; } }
[System.Diagnostics.CodeAnalysis.MemberNotNullWhen(true, "RescheduleTime")]
public bool IsRescheduleRequested { get { throw null; } }
src/api/Orleans.DurableJobs/Orleans.DurableJobs.cs:79
- The public API surface for
DurableJobRunStatusremoved thePollAfterenum member name entirely. Add it back as an[Obsolete]alias so older callers can still compile against the updated package.
public enum DurableJobRunStatus
{
Completed = 0,
InProgress = 1,
Failed = 2,
- Files reviewed: 19/19 changed files
- Comments generated: 0 new
- Review effort level: Lite
Problem
Durable job execution results distinguish completion, active polling, and failure, but cannot represent a successful execution which requests another durable run. Treating that outcome as a retry mixes success with failure policy and retry telemetry.
Solution
Clarify the serialized execution-status names and add
RescheduleAtas a successful terminal outcome for the current run. The executor routes that outcome throughIJobShard.RescheduleJobAsync, resets the persisted dequeue count, and records dedicated rescheduling logs and metrics. The in-memory and journaled shard implementations persist the reset execution generation and dequeue state. Unknown serialized status values continue through the configured failure retry policy, and the public guidance defines the rolling-upgrade order for serialized value 3.Rationale
Durable Jobs is pre-release, so the shard contract directly represents both failure retries and successful rescheduling. This preserves their distinct semantics and gives each successful rescheduled execution a fresh run identity and first-attempt dequeue count.
Microsoft Reviewers: Open in CodeFlow