Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 12 additions & 4 deletions docs/site/src/content/docs/grains/journaling/azure-storage.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,8 +19,8 @@ Configure <xref:Orleans.Journaling.AzureBlobStorageHostingExtensions.AddAzureBlo

Each journal uses:

- An append blob at `<journalId>/wal` by default.
- Immutable checkpoint blobs at `<journalId>/chk.<snapshotId>`.
- An append blob at `wal/<journalId>` by default.
- Immutable checkpoint blobs at `checkpoints/<journalId>/<snapshotId>`.
- WAL metadata which identifies the current checkpoint, journal format, and optimistic-concurrency state.

Recovery reads the published checkpoint followed by the WAL tail. A replacement uploads the new checkpoint and then atomically publishes it through WAL metadata. The provider performs best-effort cleanup of obsolete checkpoints after publication when <xref:Orleans.Journaling.AzureBlobJournalStorageOptions.DeleteOldCheckpoints> is `true`, which is the default.
Expand All @@ -29,7 +29,11 @@ Customize <xref:Orleans.Journaling.AzureBlobJournalStorageOptions.GetWalBlobName

Azure append blobs limit append-block size and block count. The provider accepts an encoded append batch up to 100 MiB, requests compaction after 49,000 committed blocks, and reserves additional headroom before the 50,000-block service limit.

The journal catalog discovers the default `/wal` naming shape. Preserve that suffix when custom names need catalog listing, or provide the application-specific discovery mechanism required by the caller.
The journal catalog selects append blobs in the `wal/` namespace of the configured container. Custom naming delegates participating in catalog discovery produce `wal/<journalId>` for each journal. The separate checkpoint namespace keeps checkpoints out of listing pages. Raw journal-id prefixes and conservative ASCII bounds narrow the native listing on both flat-namespace and hierarchical-namespace (HNS) accounts.

HNS recursive listings sort `/` before other characters. The catalog preserves each bound's shared listing prefix and widens the remaining suffix when punctuation or directory separators affect ordering. For example, an inclusive range from `a!` through `a/0` starts at `wal/a` and completes after crossing `wal/a0`. Every returned candidate is checked against the original ordinal range. Shared prefixes keep timestamp scans narrow; wider boundaries can transfer additional candidates. Both account types use the same listing algorithm and configured Blob client.

Catalog callers can request a metadata snapshot with each identity. Blob listings project the WAL's format, ETag, and caller-owned metadata in the listing response. The snapshot can replace a separate metadata read; conditional updates use its ETag to detect concurrent changes.

## Azure Table Storage

Expand All @@ -48,7 +52,11 @@ A single append batch is limited to 2 MiB by the provider's entity group transac

Compaction is requested at either <xref:Orleans.Journaling.AzureTableJournalStorageOptions.CompactionRowCountThreshold> (10,000 rows by default) or <xref:Orleans.Journaling.AzureTableJournalStorageOptions.CompactionSizeThreshold> (32 MiB by default).

Customize <xref:Orleans.Journaling.AzureTableJournalStorageOptions.GetPartitionKey> when a different partition layout is required. The mapping must remain unique per journal and satisfy Azure Table partition-key constraints.
The default partition mapping accepts printable ASCII journal ids (`0x20` through `0x7E`) and encodes each byte as two uppercase hexadecimal digits. It preserves ordinal ordering and prefixes for indexed catalog queries and supports journal ids up to 512 characters. Storage creation validates the id before accessing Azure. This restriction is specific to the default Table mapping.

Customize <xref:Orleans.Journaling.AzureTableJournalStorageOptions.GetPartitionKey> when a different partition layout is required. Custom mappings support other journal-id alphabets, remain unique per journal, and satisfy Azure Table partition-key constraints. Their catalog queries filter the canonical journal-id property and can require a table scan.

Metadata-enabled catalog queries select the header's format and caller-owned metadata together with its ETag. Identity-only queries retain their smaller projection. Each returned metadata snapshot has the same meaning as a direct metadata read and can become stale after it is listed.

## Optimistic concurrency

Expand Down
38 changes: 36 additions & 2 deletions src/AWS/Orleans.Journaling.S3/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,11 +10,45 @@ Buckets should be created ahead of time for AWS S3 Express One Zone. `CreateBuck

Metadata updates rewrite the current WAL using a conditional single-object upload. Publish a checkpoint to compact the WAL before updating metadata when the replacement object would exceed S3's 5 GB (5,000,000,000 byte) single-upload limit. Checkpoint snapshots use the same upload limit.

## Object layout

`GetObjectKey` maps a logical journal id to its base object key (the identity mapping by default). WAL and checkpoint objects use separate namespaces:

- WAL: `wal/<base-key>`
- Checkpoint: `checkpoints/<base-key>/<snapshot-id>`

For example, journal `jobs/00001234` uses `wal/jobs/00001234` and checkpoint objects under `checkpoints/jobs/00001234/`. Checkpoint names are stored in the WAL metadata. Catalog requests always stay under `wal/`, so checkpoints do not consume listing pages.

S3 Express directory buckets benefit from slash-delimited prefixes. Unordered listings scan the selected WAL directory and retain every matching overdue journal, however old. Applications can supply hierarchical base keys, provided their prefix and reverse mappings satisfy the catalog contract.

**Alpha layout upgrade:** When upgrading from the previous `<base-key>/wal` layout, drain durable jobs before deploying the new version, then recreate journals using the new layout.

## Catalog enumeration

`IJournalStorageCatalog.ListAsync` returns journal identities incrementally in S3 traversal order, including unordered directory-bucket listings. Set `ListOptions.Prefix` to select an exact journal id and its descendants.
`IJournalStorageCatalog.ListAsync` returns `JournalCatalogEntry` values incrementally in S3 traversal order. Each entry's `Id` is the journal identity. S3 entries always have null `Metadata`, including when `ListOptions.IncludeMetadata = true`: `ListObjectsV2` cannot project the complete journal format, ETag, and caller-owned properties together. Enumeration never adds separate per-journal metadata requests. Call `GetMetadataAsync` explicitly when metadata is needed.

`ListOptions.Prefix` is a raw ordinal string prefix, including partial segments; use a trailing `/` to select only entries inside a namespace. `MinId` and `MaxId` provide inclusive ordinal bounds, unlimited by default. All constraints are snapshotted when enumeration begins and checked before yielding. Empty intersections issue no request. Consumers needing due order must sort selected ids using `StringComparer.Ordinal`.

The provider handles `ListObjectsV2` continuations internally and requests up to 1000 objects per page. `UseOrderedListing` defaults to `false`, matching S3 Express directory buckets. Every native prefix begins with `wal/`, including unprefixed catalog requests. Directory mode widens a raw native prefix to its nearest slash-terminated directory boundary, retaining at least `wal/`. Directory buckets are unordered and do not support `StartAfter`: the provider scans that selected namespace and filters both bounds without stopping at the first future id.

Set `UseOrderedListing = true` only for general-purpose buckets or compatible services guaranteeing lexically ordered `ListObjectsV2` results and `StartAfter` support. With the default identity `GetObjectKey` mapping, ordered mode sends `wal/` plus the raw prefix, further narrowed by the common prefix of both bounds when possible. Since `StartAfter` is exclusive, an ASCII lower bound uses a strictly earlier marker: `wal/` plus the lower bound with its final character removed. This marker is sent only when it sorts after the native prefix, and the inclusive minimum is still checked locally. Subsequent requests use the returned opaque `ContinuationToken` alone to resume traversal, including after empty pages. The provider stops after `wal/<MaxId>` for a safe ASCII maximum. Non-ASCII bounds are filtered without unsafe native key seeks or cutoffs. There is no native upper-end parameter, so the crossing page is fetched but later pages are not. `UseS3ExpressAppend` does not imply a listing-order guarantee.

Custom `GetObjectKey` mappings must also configure `GetObjectKeyPrefix` for explicitly prefixed catalog listings. Arbitrary object-key mapping functions cannot safely be applied to a journal prefix. The explicit prefix mapper must return a non-empty **base object-key prefix** covering every matching journal, including partial segments; do not include `wal/`. The provider prepends `wal/`, then directory mode widens the result to a slash boundary. Missing or empty prefix configuration throws before listing; unprefixed listing, including bounds-only queries, still works without the mapper. Custom mappings never use identity-based native lower or upper key bounds, even in ordered mode.

```csharp
options.GetObjectKey = id => $"journals/{id.Value}";
options.GetObjectKeyPrefix = prefix => $"journals/{prefix.Value}";
options.TryParseJournalId = key => key.StartsWith("journals/", StringComparison.Ordinal)
? new JournalId(key["journals/".Length..]) : null;
```

`TryParseJournalId` receives the base object key after the catalog strips `wal/`. Canonical-WAL validation still applies after parsing. Checkpoints are outside the selected namespace; aliases and unrelated objects under `wal/` still consume space in the native page before filtering.

The provider handles `ListObjectsV2` continuations internally, requests up to 1000 objects per page, and yields canonical WAL identities from that page before fetching more objects. The bucket traversal supports `GetObjectKey` and `TryParseJournalId` mappings; checkpoints, aliases, and unrelated objects consume space in the native page before filtering.
| Listing mode | Native prefix | Lower bound | Upper bound |
| --- | --- | --- | --- |
| Ordered, default identity mapping | `wal/` + raw prefix/common range prefix | Strictly earlier ASCII `StartAfter` marker when it narrows the prefix | Stop after crossing ASCII `wal/<MaxId>` |
| Directory/unordered | Nearest slash-terminated directory under `wal/` | Local filtering | Local filtering; entire selected namespace is traversed |
| Custom mapping | `wal/` + explicit mapped base prefix, widened in directory mode | Local filtering | Local filtering; no identity-order assumption |

Client traversal memory is proportional to the current native page. An enumerator advance can cross multiple filtered or empty pages, and the storage service determines scan work, latency, and retries. Enumeration observes the live bucket; concurrent changes follow S3 listing semantics. Use subsequent enumerations to discover later changes and tolerate repeated identities during changes.

Expand Down
68 changes: 65 additions & 3 deletions src/AWS/Orleans.Journaling.S3/S3JournalStorageOptions.cs
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,8 @@ namespace Orleans.Journaling;
/// </summary>
public sealed class S3JournalStorageOptions
{
internal const string WalObjectKeyPrefix = "wal/";

private IAmazonS3? _s3Client;

/// <summary>
Expand All @@ -20,11 +22,45 @@ public sealed class S3JournalStorageOptions
/// <summary>
/// Gets or sets the delegate used to generate the base object key for a journal.
/// </summary>
/// <remarks>
/// WAL objects use <c>wal/&lt;base-key&gt;</c> and checkpoint objects use
/// <c>checkpoints/&lt;base-key&gt;/&lt;snapshot-id&gt;</c>. The delegate must not add these namespaces.
/// </remarks>
public Func<JournalId, string> GetObjectKey { get; set; } = DefaultGetObjectKey;

/// <summary>
/// Gets or sets the delegate used to parse journal ids from catalog object keys.
/// Gets or sets the delegate mapping a non-default catalog prefix to a base object-key prefix.
/// </summary>
/// <remarks>
/// The result must be non-empty and include the base object key of every journal whose id
/// starts with the supplied raw ordinal prefix, including partial segments. Additional matches
/// are filtered by the catalog. When unset, the default identity <see cref="GetObjectKey"/>
/// mapping uses the raw journal prefix. Custom object-key mappings must configure this delegate
/// to use prefixed listings. Unprefixed listings do not require this delegate.
/// The provider prepends <c>wal/</c> to the result; the delegate must not add it.
/// When <see cref="UseOrderedListing"/> is false, the mapped prefix is widened to its nearest
/// slash-terminated directory boundary to support directory buckets, while retaining <c>wal/</c>.
/// </remarks>
public Func<JournalId, string>? GetObjectKeyPrefix { get; set; }

/// <summary>
/// Gets or sets a value indicating that the bucket supports lexically ordered
/// <see cref="Amazon.S3.Model.ListObjectsV2Request"/> listings. Defaults to false.
/// </summary>
/// <remarks>
/// Enable this only for general-purpose buckets or compatible services which guarantee ordered
/// listings and support <see cref="Amazon.S3.Model.ListObjectsV2Request.StartAfter"/>.
/// Directory buckets are unordered and must leave this disabled. With the default identity
/// object-key mapping, ordered listings can seek to an ASCII lower bound and stop beyond an
/// ASCII upper bound. Custom mappings do not use native key-range bounds.
/// This capability is independent of <see cref="UseS3ExpressAppend"/>.
/// </remarks>
public bool UseOrderedListing { get; set; }

/// <summary>
/// Gets or sets the delegate used to parse journal ids from base object keys.
/// </summary>
/// <remarks>The catalog removes the <c>wal/</c> namespace before invoking this delegate.</remarks>
public Func<string, JournalId?> TryParseJournalId { get; set; } = DefaultTryParseJournalId;

/// <summary>
Expand Down Expand Up @@ -122,6 +158,32 @@ public IAmazonS3? S3Client

internal bool IsClientExternallyOwned { get; private set; }

internal bool UsesDefaultObjectKey => GetObjectKey == DefaultGetObjectKey;

internal string? GetObjectKeyPrefixForCatalog(string? prefix)
{
if (prefix is null)
{
return null;
}

var mapper = GetObjectKeyPrefix;
if (mapper is null && !UsesDefaultObjectKey)
{
throw new InvalidOperationException(
$"Custom {nameof(GetObjectKey)} mappings require {nameof(GetObjectKeyPrefix)} for prefixed journal listings.");
}

var result = mapper is null ? prefix : mapper(new JournalId(prefix));
if (string.IsNullOrEmpty(result))
{
throw new InvalidOperationException(
$"{nameof(GetObjectKeyPrefix)} must return a non-empty object-key prefix.");
}

return result;
}

internal string GetObjectKeyForJournal(JournalId journalId)
{
if (journalId.IsDefault)
Expand All @@ -147,9 +209,9 @@ internal static string GetCheckpointObjectKeyForJournal(JournalId journalId, str
return GetDefaultCheckpointObjectKey(journalObjectKey, snapshotId);
}

internal static string GetDefaultWalObjectKey(string journalObjectKey) => $"{journalObjectKey}/wal";
internal static string GetDefaultWalObjectKey(string journalObjectKey) => $"{WalObjectKeyPrefix}{journalObjectKey}";

internal static string GetDefaultCheckpointObjectKey(string journalObjectKey, string snapshotId) => $"{journalObjectKey}/chk.{snapshotId}";
internal static string GetDefaultCheckpointObjectKey(string journalObjectKey, string snapshotId) => $"checkpoints/{journalObjectKey}/{snapshotId}";

internal Func<CancellationToken, Task<IAmazonS3>> GetCreateClient()
=> CreateClient ?? (_ => Task.FromResult<IAmazonS3>(new AmazonS3Client(ClientConfig ?? new AmazonS3Config())));
Expand Down
Loading
Loading