Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 3 additions & 2 deletions docs/site/scripts/lib/link-audit.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -757,6 +757,8 @@ export async function auditRenderedInternalLinks({
targetUrl.origin !== 'https://dotnet.github.io' &&
targetUrl.origin !== 'https://docs.invalid'
) {
const normalized = new URL(targetUrl);
normalized.hash = '';
const sourceTarget = repositorySourceTarget(targetUrl, repositoryRoot);
if (sourceTarget) {
if (sourceTarget.error) {
Expand All @@ -779,11 +781,10 @@ export async function auditRenderedInternalLinks({
);
}
}
externalTargets?.delete(normalized.href);
continue;
}
if (['http:', 'https:'].includes(targetUrl.protocol) && externalTargets) {
const normalized = new URL(targetUrl);
normalized.hash = '';
if (!externalTargets.has(normalized.href)) {
externalTargets.set(normalized.href, [
{ relativeFile: route, line: 1, rendered: true },
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,14 @@ options, storage, catalog, and state-manager factory share the same binding.
The unnamed overload configures default grain journaling; named registrations
leave that default independent.

For an account, container, table, or backend cutover, retain the old namespace
and its mapping under its original binding while new Durable Jobs shards use
a second binding. During draining, the old account still needs listing,
conditional ownership updates, journal mutations, and delete permissions.
Granting only read access prevents successful draining. See
[Migrate Durable Jobs storage](durable-jobs-migration.md) for the staged rollout,
full inventory, and retirement criteria.

## Azure Blob Storage

Configure <xref:Orleans.Journaling.AzureBlobStorageHostingExtensions.AddAzureBlobJournalStorage*> with an authenticated <xref:Azure.Storage.Blobs.BlobServiceClient>:
Expand Down
16 changes: 15 additions & 1 deletion docs/site/src/content/docs/grains/journaling/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,20 @@ it contains work. A configuration-based `GrainJournaling` provider name selects
the journal registration; `ServiceKey` selects the Azure or Redis client from
dependency injection.

For Durable Jobs, <xref:Orleans.Hosting.DurableJobsExtensions.UseJournaledDurableJobs*>
selects the journaled implementation.
<xref:Orleans.Hosting.DurableJobsOptions.ActiveProviderName> selects the provider
for new shards and defaults to `Default`.
<xref:Orleans.Hosting.DurableJobsOptions.DrainingProviderNames> selects additional
providers containing existing shards and is empty by default. All selected
bindings require storage, catalog, and factory services and are validated at
startup. Bindings remain fixed for that process's lifetime.

See [Migrate Durable Jobs storage](durable-jobs-migration.md) for named selection,
deployment staging, and full-namespace retirement checks. The runnable sample
linked from that guide exercises the new APIs against packages built from this
repository; existing published snippet packages predate named-provider selection.

## Configure the JSON format

JSON Lines is the default write format. Register source-generated metadata for every application type used as a durable key, value, collection item, persistent state, or durable task result:
Expand Down Expand Up @@ -88,6 +102,6 @@ The default minimum is seven days. Removal is persisted by a compaction after th

## Development storage

<xref:Orleans.Journaling.HostingExtensions.AddJournalStorage*> registers core services and resolves an <xref:Orleans.Journaling.IJournalStorageProvider>. Runtime tests and disposable development hosts can use <xref:Orleans.Journaling.HostingExtensions.AddVolatileJournalStorage*> with a provider name. Its contents live in process memory, so use persistent emulator storage to validate restart recovery.
<xref:Orleans.Journaling.HostingExtensions.AddJournalStorage*> registers core services and resolves an <xref:Orleans.Journaling.IJournalStorageProvider>. Runtime tests and disposable development hosts can use <xref:Orleans.Journaling.HostingExtensions.AddVolatileJournalStorage*> with a provider name. Its contents live in process memory, so use persistent emulator storage to validate restart recovery and provider migration.

Use the same durable provider category in staging that production uses so recovery, compaction, concurrency, and backup procedures receive realistic validation.
Original file line number Diff line number Diff line change
@@ -0,0 +1,191 @@
---
title: Migrate Durable Jobs storage
description: Move new Durable Jobs to a named journal provider while draining existing storage and verifying retirement.
ms.date: 09/16/2026
ms.topic: how-to
---

# Migrate Durable Jobs storage

Durable Jobs uses journal storage for shard metadata and job records. Select a
write provider for new shards and keep previous providers selected for draining
their existing jobs. Most deployments use a single provider; additional
providers are needed only while work remains in other storage namespaces.

This procedure directs **new work** to B while existing journals drain in place.
Keep one authoritative location for each shard throughout its lifetime.

## Prepare the bindings

1. Register A's existing namespace and B's new namespace using the named
[Azure Blob/Table](azure-storage.md#named-storage-namespaces),
[Redis](redis-journal-storage.md#configure-redis-journal-storage), or S3
journal registration. Use separate accounts, containers, tables, buckets,
or key prefixes as appropriate. Each selected namespace has one binding.
1. Retain A's identity, namespace mapping, journal format readers, and payload
codecs. Its credentials must permit reads, listing, conditional metadata
updates, journal writes, compaction, and deletion throughout the drain.
1. Register <xref:Orleans.Hosting.DurableJobsExtensions.UseJournaledDurableJobs*>
and configure <xref:Orleans.Hosting.DurableJobsOptions.ActiveProviderName>
and <xref:Orleans.Hosting.DurableJobsOptions.DrainingProviderNames>.
The default write name is `Default`; the draining list is initially empty.
1. Validate both namespaces and permissions in an isolated environment. Startup
validates selected storage, catalog, and factory bindings. Physical
availability and access still need operational checks.

Named journal providers resolve storage, catalog, and state-manager factory
from the same binding. The default provider used by
<xref:Orleans.Journaling.DurableGrain> remains independent of explicitly named
Durable Jobs providers. The in-memory and Azure Durable Jobs convenience
methods compose the same common registration path for the default binding.

For executable configuration and verification, use the
[Durable Jobs migration sample](https://github.com/dotnet/orleans/tree/main/samples/DurableJobsMigration).
It runs separate prepare/drain console processes against disk-backed Azurite,
persists a future job handle, recovers and executes A work after restart, creates
B work, and reports complete inventories locally. Named-provider APIs require
packages built from sources containing this feature; the repository's sample
build supplies those packages.

## Stage the deployment

Provider selection is fixed at process startup. Restart silos when changing
bindings or selection.

| Stage | Write provider | Draining providers | Required condition |
| --- | --- | --- | --- |
| Steady state | A | Empty | All work resides in A |
| Prepare | A | B | Every scheduling silo can discover and locate both A and B before B receives new work |
| Cut over | B | A | Restart all scheduling silos with this selection and finish prior scheduling calls |
| Retire | B | Empty | Verify the complete drain and stop every remaining A writer first |

During an ordinary rolling cutover, older A-configured silos can still create
A shards. Treat deployment completion and the end of admitted scheduling calls
as part of the boundary. If the application requires a strict cutover time,
pause application scheduling, finish in-flight scheduling, update all scheduling
silos, verify their configuration, and then resume.

Retain A's credentials and namespace through its final cleanup. Existing A work
continues mutating A while B receives new jobs.

## Understand routing during the drain

New schedules use newly created shards in B. Discovery sweeps A and B with one
oldest-first ordering and aggregate claim budget, then opens each shard with
the binding of the provider which contained it. Execution, ownership changes,
retries, rescheduling, cancellation, snapshots, and deletion retain that binding.
Draining shards stay closed to new schedules.

Shard IDs retain their timestamp/GUID representation, and journal IDs remain
`jobs/shards/<id>`. Serialized <xref:Orleans.DurableJobs.DurableJob> handles keep
their existing provider-independent shape.
Application code can keep using a previously saved handle after cutover.

For an uncached known shard:

- With one selected provider, the runtime reads that provider's metadata
directly. This path performs one metadata read regardless of other registered
providers.
- With multiple selected providers, it reads the exact journal ID in the write
provider first, then draining providers as needed. This lookup also finds
future shards outside normal discovery lookahead.
- Once the shard is located, operations use its original binding. A tracked
shard already has this binding.

Successful absence checks across all selected providers establish absence.
Provider failures that leave a lookup unresolved surface as an
<xref:System.AggregateException>. If another provider supplies the authoritative
shard, that binding can be used while the earlier failure remains observable.
Discovery reports each provider's failures with its configured name. A later
fresh sweep retries failed providers.

Each sweep enumerates the selected catalogs in configured order and buffers each
provider's candidates until its enumeration succeeds. Shard assignment starts
after all selected catalogs have completed or faulted, using the successful
results for the shared oldest-first claim budget. A catalog failure discards that
provider's partial results.

Recovery latency therefore includes catalog listing requests and storage-client
retry delays. Configure request timeouts and retry limits on each storage client
to match the required recovery latency, and monitor catalog latency during the
drain. The shard check interval schedules fresh sweeps; storage-client settings
govern pending request attempts.

## Verify retirement

Resolve <xref:Orleans.DurableJobs.IDurableJobsStorageInspector> from the host's
services and call its <xref:Orleans.DurableJobs.IDurableJobsStorageInspector.InspectAsync*>
method with A's configured name and a cancellation token.
The returned <xref:Orleans.DurableJobs.DurableJobsStorageStatus> describes a full
`jobs/shards/` catalog snapshot through the caller using read-only catalog and
metadata operations. It includes future shards outside lookahead and poisoned
shards requiring operator recovery.

| Result | Operational meaning |
| --- | --- |
| `ProviderName`, `IsWriteProvider` | The caller's configured binding and write selection |
| `ShardCount` | Existing journals in the shard namespace, including unrecognized entries; each recognized shard can contain many jobs |
| `OwnedShardCount` | Shards with an owner recorded in metadata; evaluate owner liveness separately through cluster membership |
| `PoisonedShardCount` | Shards requiring investigation before normal processing can continue |
| `UnrecognizedShardCount` | Entries with uninterpretable shard metadata; resolve them before retirement |
| `OldestShardStartTime`, `NewestShardStartTime` | Nullable bounds from recognized shard-window metadata; retries and reschedules can move attempts beyond these windows |

Counts are 64-bit values. Owned and poisoned counts can overlap. Inspection
failure or cancellation faults the operation. A successful result covers the
complete enumeration. Display **unknown** and preserve the error on failure.

Require all of the following before removing A:

- Every scheduling silo and other A writer has cut over, and prior scheduling
requests have finished.
- A successful complete inventory reports a total shard count of zero,
including unrecognized entries.
- Future-dated, owned, poisoned, and repeatedly rescheduled work has been
resolved through the application's recovery/cancellation procedures.
- Completion and empty-shard deletion have succeeded; storage authorization and
cleanup errors have been investigated.
- Repeated inventories, as required by the backend's live-listing semantics,
agree with the deployment evidence.
- Other consumers of the same binding have finished using it.

Retirement requires both a successful full-zero inventory and completed
cluster-wide writer cutover. An inventory is a live observation, and
`IsWriteProvider` reflects the inspecting process's configuration. Keep
independent deployment evidence for every writer's cutover and the time of
each successful inventory.

Remove A from the draining list and restart with B alone after these conditions
hold. Retire A's storage binding and credentials only when all other consumers
are also finished. The runtime returns to the single-provider direct-lookup
path. Handles for already completed jobs retain the existing cancellation
outcome when their shards are absent.

## Roll back safely

Keep both providers selected and retain provider-aware binaries. Switch writes
back to A while B drains, using the same staged-deployment discipline. Existing
B shards keep executing and mutating B. Keep both namespaces, permissions,
format readers, and payload codecs available through the rollback window.

## Monitor and diagnose

Follow [provider-scoped drain monitoring](operations.md#monitor-durable-jobs-provider-draining).
Record inventory timestamps and failures alongside deployment versions and
configured provider names. Successful inspection emits a structured
Information-level inventory log; provider failures emit structured Error-level
logs with the configured name. Use these logs and inspector snapshots for
provider-scoped progress, alongside existing Durable Jobs and journal metrics.

| Symptom | Check and response |
| --- | --- |
| A's count grows after B is enabled | Find silos or other writers still configured to create A shards; complete cutover before evaluating retirement |
| A remains populated while due-job execution is idle | Inspect the full namespace for future, poisoned, owned, or repeatedly rescheduled work |
| A jobs execute but its count remains above zero | Investigate remaining writable shards, retry/reschedule activity, and failed empty-shard deletion |
| Cancellation after cutover reports a storage error | Check reachability and mutation permissions of every selected provider; preserve the storage error and restore access before retrying the request |
| One provider's inventory is stale | Mark its status unknown; inspect listing, metadata reads, throttling, and authorization |

Inventory can be expensive on large catalogs. Run it at a controlled cadence
with a bounded operator timeout, restrict access to operational results, and
keep credentials out of logs. The sample reports inventory through local
console output. Apply authentication and authorization to any operational
interfaces added by a deployment.
34 changes: 34 additions & 0 deletions docs/site/src/content/docs/grains/journaling/operations.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,40 @@ At minimum, dashboard and alert on:

Correlate these signals with grain identity, provider dependency health, deployment version, and storage throttling.

### Monitor Durable Jobs provider draining

Track the configured provider name separately from the backend category in
journaling catalog metrics. Two named Blob providers use the same backend but
can have different availability and remaining work.

Use <xref:Orleans.DurableJobs.IDurableJobsStorageInspector.InspectAsync*> to
collect a complete shard-namespace snapshot per selected provider. Record the
inventory's start/completion timestamps, caller, configured write provider,
success/failure, counts, and oldest/newest shard start time. Keep the last
successful result with its timestamp for history; mark current status unknown
on failure.

Successful inspection also emits a structured Information-level inventory log.
Provider errors emit structured Error-level logs with the configured provider
name. Use inspector snapshots and these logs for provider-scoped drain progress;
existing execution and catalog metrics provide complementary workload signals.

During cutover, alert on old-provider write-permission failures, failed
discovery or known-ID lookups, inventory failures, poisoned or unrecognized
entries, and an unexpectedly flat drain curve. Each shard may contain many jobs,
and retries or rescheduling can extend its
lifetime beyond its original time window. Oldest/newest shard start times
describe recognized shard windows; retries and reschedules can move attempts
beyond those windows.

Normal discovery uses lookahead bounds and can skip future or poisoned work.
Use the complete inspector results and deployment evidence in the
[retirement checklist](durable-jobs-migration.md#verify-retirement) to decide
whether old storage can be removed. Execution throughput and eligible-work
sweeps provide complementary workload signals. Run inspection as a protected
operator command or background diagnostic task, and apply authentication and
authorization to operational interfaces.

### Monitor catalog traversal and retries

Use `orleans-journaling-provider-catalog-pages` to track pages received during S3 and Azure catalog
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,14 @@ configure default grain journaling. A `ServiceKey` in configuration chooses the
client connection, while the `GrainJournaling` provider name chooses the storage
binding.

Durable Jobs can select a Redis write provider and retain another provider for
draining. Keep the old prefix, credentials, and mutation permissions until every
old shard is resolved: claims, retries, cancellation, compaction, and deletion
continue there. A Redis catalog scan has live-listing semantics, so successful
inventory is an observation rather than a transactionally frozen cluster view.
Follow [Migrate Durable Jobs storage](durable-jobs-migration.md), including
repeated complete inventories after all scheduling silos have cut over.

## Storage behavior

The provider stores journal data in Redis strings and journal metadata in Redis hashes. Per-journal reads and mutations use atomic Lua scripts. Catalog operations discover journals by scanning metadata keys on each connected primary Redis server.
Expand Down
2 changes: 2 additions & 0 deletions docs/site/src/content/docs/toc.yml
Original file line number Diff line number Diff line change
Expand Up @@ -98,6 +98,8 @@ items:
items:
- name: Overview
href: grains/journaling/index.md
- name: Migrate Durable Jobs storage
href: grains/journaling/durable-jobs-migration.md
- name: Use durable state
href: grains/journaling/durable-state.md
- name: Runtime behavior and consistency
Expand Down
13 changes: 11 additions & 2 deletions docs/site/tests/link-audit.test.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -392,7 +392,7 @@ describe('rendered internal link audit', () => {
);
});

test('validates generated repository source links locally', async () => {
test('validates source and generated repository links locally', async () => {
const repositoryRoot = await temporaryDirectory();
const distRoot = path.join(repositoryRoot, 'dist');
await mkdir(path.join(repositoryRoot, 'src'));
Expand All @@ -413,7 +413,16 @@ describe('rendered internal link audit', () => {
'<a href="https://github.com/dotnet/orleans/pulls">non-source repository URL</a>',
].join(''),
);
const externalTargets = new Map();
const externalTargets = new Map([
[
`https://github.com/dotnet/orleans/blob/${commit}/src/Widget.cs`,
[{ relativeFile: 'guide.md', line: 1 }],
],
[
'https://github.com/dotnet/orleans/tree/main/samples/Example',
[{ relativeFile: 'guide.md', line: 2 }],
],
]);

const issues = await auditRenderedInternalLinks({
distRoot,
Expand Down
Loading
Loading