fix(dr): make backups restorable, reliable, complete, and alerting - #986
Conversation
D1's import path rejects statements above its ~100 KB limit (SQLITE_TOOBIG) and exports write one INSERT per row, so a single oversized row made every production backup un-importable (verified with live restore drills of the 2026-07-26 export). - add shared restore-safety limits (packages/shared/src/backup-restore-safety.ts) - drop oversized package-invocation replay caches instead of storing them (duplicates get the existing idempotency_response_unavailable outcome) - truncate stored email body copies at 64 KiB (raw MIME in R2 stays canonical) - reject oversized value_set writes with a storage-bucket hint - migration 0102 bounds existing rows (81 oversized invocation rows in production, 2 oversized email bodies) Co-authored-by: Kent C. Dodds <me+github@kentcdodds.com>
The finalize step re-polled D1 with the cached bookmark and required the fresh download to byte-match the stored object. When the short-lived poll result had expired, the refresh started a new export of a newer database state, so the comparison was doomed whenever production wrote anything in between — roughly every other nightly backup errored terminally with existing-object-source-mismatch and left orphaned ~107 MB immutable objects behind. Finalization now verifies the stored object against the durable upload-step digest (size, R2 ETag, full SHA-256 re-read) and never polls D1 again. It also measures statement lengths while streaming (quote-aware) and persists <objectKey>.stats.json beside the SQL; oversized statements log backup-unrestorable-statements with failure status because such a backup cannot be re-imported through the D1 API. Also refreshes TRUSTED_RESTORE_BASELINE_SHA256 (stale since #904) for the current migration set including 0102. Co-authored-by: Kent C. Dodds <me+github@kentcdodds.com>
The staging exporter's 00:30-02:10 UTC window was never enough at production scale: every night ended mid-artifacts-phase, so exporter/summary.json was never written and no day could ever be sealed. Extend the window to 00:30-06:10 (completed days exit via the cheap already-complete check) and add a 06:15 watchdog lane that fails loudly to Sentry when the summary is still missing — window exhaustion was previously silent. Documents the new schedule, the restore-safe row size contract, and the TRUSTED_RESTORE_BASELINE_SHA256 recipe in the disaster-recovery runbook. Co-authored-by: Kent C. Dodds <me+github@kentcdodds.com>
|
Warning Review limit reached
Next review available in: 56 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (24)
📝 WalkthroughWalkthroughThe change adds restore-safe UTF-8 and database value limits, records SQL statement statistics during backups, revises immutable-manifest finalization, adds a staging export watchdog, updates scheduled execution, and documents the resulting recovery and alerting behavior. ChangesDisaster recovery safeguards
Estimated code review effort: 4 (Complex) | ~45 minutes Sequence Diagram(s)sequenceDiagram
participant BackupRuntime
participant ImmutableStorage
participant BackupBucket
participant Manifest
BackupRuntime->>ImmutableStorage: stream and digest SQL object
ImmutableStorage-->>BackupRuntime: integrity data and SQL statement stats
BackupRuntime->>Manifest: write immutable manifest
BackupRuntime->>BackupBucket: write object-adjacent stats JSON
sequenceDiagram
participant WorkerScheduler
participant DrExportWatchdog
participant StagingBucket
WorkerScheduler->>DrExportWatchdog: invoke watchdog tick
DrExportWatchdog->>StagingBucket: read exporter/summary.json
StagingBucket-->>DrExportWatchdog: summary present or missing
DrExportWatchdog-->>WorkerScheduler: return result or throw progress error
Possibly related PRs
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
main gained 0102-0104 while this branch was open; the data migration is content-identical, renumbered to the next free prefix with the ledger entry and TRUSTED_RESTORE_BASELINE_SHA256 recomputed. Co-authored-by: Kent C. Dodds <me+github@kentcdodds.com>
Co-authored-by: Kent C. Dodds <me+github@kentcdodds.com>
|
🔎 Preview deployed: https://kody-pr-986.kody-a99.workers.dev Worker: Mocks:
|
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit b8a4398. Configure here.
| text_body: input.message.textBody ?? null, | ||
| html_body: input.message.htmlBody ?? null, | ||
| text_body: boundedEmailBody(input.message.textBody), | ||
| html_body: boundedEmailBody(input.message.htmlBody), |
There was a problem hiding this comment.
Dual email bodies break limit
Medium Severity
boundedEmailBody caps text_body and html_body independently at maxRestorableTextColumnBytes, but D1 exports one INSERT per row and the import limit applies to the entire statement. A message with both bodies near the cap can still produce an INSERT well above d1ImportMaxStatementBytes, leaving backups unrestorable despite the new safety checks.
Reviewed by Cursor Bugbot for commit b8a4398. Configure here.
There was a problem hiding this comment.
Actionable comments posted: 3
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
packages/backup-control-plane/immutable-storage.ts (1)
93-131: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick winCancel the abandoned digest reader when
digestBodyfails.In
storeSignedDownload, the download body is teed for concurrent R2 upload and digesting. IfdigestBodyrejects and only callswriter.abort(), thedigestBodyStreamreader is left unreleased whilebucket.put()continues draining the shared source. Explicitly cancel that branch on error so the source/backpressure can be released cleanly.🔧 Proposed fix
} catch (error) { await writer.abort(error).catch(() => undefined) + await reader.cancel(error).catch(() => undefined) throw error }🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@packages/backup-control-plane/immutable-storage.ts` around lines 93 - 131, Update the error path in digestBody to cancel the digest body reader before or alongside aborting the writer. Ensure reader.cancel(error) is awaited or safely handled, while preserving the existing writer.abort(error) cleanup and rethrow behavior so the abandoned tee branch releases its source cleanly.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@packages/backup-control-plane/backup-runtime.ts`:
- Around line 230-241: Make the advisory record-statement-stats step non-fatal
after the immutable manifest is successfully written: isolate the await of
recordSqlStatementStats from the outer backup failure path, catch exhausted
retries, and log the stats error without rethrowing or emitting backup-failure.
Preserve the existing retries and timeout while allowing the overall backup
operation to remain successful when stats persistence fails.
In `@packages/shared/src/backup-restore-safety.ts`:
- Around line 20-25: Replace the universal raw-column cap in
packages/shared/src/backup-restore-safety.ts:20-25 with a serialized SQL-row
budget that accounts for apostrophe escaping and table-specific co-resident
fields. Apply that shared budget in
packages/worker/migrations/0102-restore-safe-row-sizes.sql:12-31, including
combined text_body/html_body limits; update
packages/worker/src/package-invocations/repo.ts:21-33 and
packages/worker/src/mcp/values/service.ts:69-75 to validate escaped SQL-storage
size. Extend
packages/worker/src/app/restore-safe-row-sizes-migration.node.test.ts:75-109 and
packages/worker/src/mcp/values/service.node.test.ts:377-406 with worst-case
apostrophe-heavy and dual-large-body boundary cases.
In `@packages/worker/src/email/repo.ts`:
- Around line 33-40: Replace the independent per-column truncation in
boundedEmailBody with a shared encoded-row budget allocated across both
text_body and html_body. Ensure the combined UTF-8 byte lengths, truncation
notices, and both stored values remain within the aggregate limit, while
preserving null handling and valid UTF-8. Update the email import coverage to
include both bodies near their individual limits and verify the resulting
combined row stays within budget.
---
Outside diff comments:
In `@packages/backup-control-plane/immutable-storage.ts`:
- Around line 93-131: Update the error path in digestBody to cancel the digest
body reader before or alongside aborting the writer. Ensure reader.cancel(error)
is awaited or safely handled, while preserving the existing writer.abort(error)
cleanup and rethrow behavior so the abandoned tee branch releases its source
cleanly.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: a3999774-056b-4540-9f71-6587853acd76
📒 Files selected for processing (24)
docs/contributing/disaster-recovery.mdpackages/backup-control-plane/backup-control-plane-test-support.tspackages/backup-control-plane/backup-runtime.node.test.tspackages/backup-control-plane/backup-runtime.tspackages/backup-control-plane/backup-types.tspackages/backup-control-plane/immutable-storage.node.test.tspackages/backup-control-plane/immutable-storage.tspackages/backup-control-plane/readme.mdpackages/backup-control-plane/wrangler.jsoncpackages/shared/src/backup-restore-safety.node.test.tspackages/shared/src/backup-restore-safety.tspackages/worker/migrations/0102-restore-safe-row-sizes.sqlpackages/worker/src/app/restore-safe-row-sizes-migration.node.test.tspackages/worker/src/dr/exporter.node.test.tspackages/worker/src/dr/exporter.tspackages/worker/src/email/repo-body-limits.workers.test.tspackages/worker/src/email/repo.tspackages/worker/src/index.tspackages/worker/src/index.workers.test.tspackages/worker/src/mcp/values/service.node.test.tspackages/worker/src/mcp/values/service.tspackages/worker/src/package-invocations/repo.tspackages/worker/src/package-invocations/service.node.test.tstools/migration-ledger.json
| await step.do( | ||
| 'record-statement-stats', | ||
| { retries: { limit: 2, delay: '10 seconds' }, timeout: '2 minutes' }, | ||
| async () => | ||
| recordSqlStatementStats({ | ||
| env, | ||
| day: checkedPayload.day, | ||
| instanceId: event.instanceId, | ||
| objectKey: stored.objectKey, | ||
| stats: stored.sqlStatementStats, | ||
| }), | ||
| ) |
There was a problem hiding this comment.
🩺 Stability & Availability | 🟠 Major | ⚡ Quick win
Advisory stats step can turn a successful backup into a reported failure.
record-statement-stats runs after the immutable manifest is already durably written, but it's awaited directly inside the same try that wraps the whole function. If it exhausts its retries (e.g. a transient R2 hiccup while writing ${objectKey}.stats.json), the error propagates to the outer catch, which logs backup-failure and rethrows — even though the backup already succeeded and is restorable. The function's own comment calls this data "advisory," which contradicts making it fail-closed for the entire run. Per the README, freshness ticks can restart an "errored" instance, so this could also trigger unnecessary restarts/alerts for a day that already has a valid manifest.
🔧 Proposed fix
- await step.do(
- 'record-statement-stats',
- { retries: { limit: 2, delay: '10 seconds' }, timeout: '2 minutes' },
- async () =>
- recordSqlStatementStats({
- env,
- day: checkedPayload.day,
- instanceId: event.instanceId,
- objectKey: stored.objectKey,
- stats: stored.sqlStatementStats,
- }),
- )
+ await step
+ .do(
+ 'record-statement-stats',
+ { retries: { limit: 2, delay: '10 seconds' }, timeout: '2 minutes' },
+ async () =>
+ recordSqlStatementStats({
+ env,
+ day: checkedPayload.day,
+ instanceId: event.instanceId,
+ objectKey: stored.objectKey,
+ stats: stored.sqlStatementStats,
+ }),
+ )
+ .catch((error) => {
+ // Advisory only: the manifest is already durable, so a failure
+ // here must not turn a successful backup into a reported failure.
+ safeLog({
+ event: 'backup-sql-stats',
+ status: 'failure',
+ day: checkedPayload.day,
+ instanceId: event.instanceId,
+ objectKey: stored.objectKey,
+ errorCode: errorCode(error),
+ })
+ })Happy to also add a regression test exercising a failing record-statement-stats step against an already-successful manifest write, if useful.
📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| await step.do( | |
| 'record-statement-stats', | |
| { retries: { limit: 2, delay: '10 seconds' }, timeout: '2 minutes' }, | |
| async () => | |
| recordSqlStatementStats({ | |
| env, | |
| day: checkedPayload.day, | |
| instanceId: event.instanceId, | |
| objectKey: stored.objectKey, | |
| stats: stored.sqlStatementStats, | |
| }), | |
| ) | |
| await step | |
| .do( | |
| 'record-statement-stats', | |
| { retries: { limit: 2, delay: '10 seconds' }, timeout: '2 minutes' }, | |
| async () => | |
| recordSqlStatementStats({ | |
| env, | |
| day: checkedPayload.day, | |
| instanceId: event.instanceId, | |
| objectKey: stored.objectKey, | |
| stats: stored.sqlStatementStats, | |
| }), | |
| ) | |
| .catch((error) => { | |
| // Advisory only: the manifest is already durable, so a failure | |
| // here must not turn a successful backup into a reported failure. | |
| safeLog({ | |
| event: 'backup-sql-stats', | |
| status: 'failure', | |
| day: checkedPayload.day, | |
| instanceId: event.instanceId, | |
| objectKey: stored.objectKey, | |
| errorCode: errorCode(error), | |
| }) | |
| }) |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/backup-control-plane/backup-runtime.ts` around lines 230 - 241, Make
the advisory record-statement-stats step non-fatal after the immutable manifest
is successfully written: isolate the await of recordSqlStatementStats from the
outer backup failure path, catch exhausted retries, and log the stats error
without rethrowing or emitting backup-failure. Preserve the existing retries and
timeout while allowing the overall backup operation to remain successful when
stats persistence fails.
| /** | ||
| * Upper bound for any single large text column persisted to D1. Leaves | ||
| * headroom below {@link d1ImportMaxStatementBytes} for the other row | ||
| * columns, INSERT framing, and quote-escaping expansion in SQL dumps. | ||
| */ | ||
| export const maxRestorableTextColumnBytes = 65_536 |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift
Raw column limits do not guarantee importable SQL statements.
A 65,536-byte value of ' expands to roughly 131 KiB when exported as a SQL string literal, before INSERT framing. Email rows can additionally contain both large body columns. This still permits SQLITE_TOOBIG, defeating the restore-safety guarantee.
packages/shared/src/backup-restore-safety.ts#L20-L25: replace the universal raw-column cap with a serialized-row budget that accounts for SQL escaping and table-specific co-resident fields.packages/worker/migrations/0102-restore-safe-row-sizes.sql#L12-L31: bound existing rows using that same budget, including the combinedtext_body/html_bodysize.packages/worker/src/app/restore-safe-row-sizes-migration.node.test.ts#L75-L109: add worst-case apostrophe and dual-large-body migration cases.packages/worker/src/package-invocations/repo.ts#L21-L33: reject/null responses based on their escaped SQL-storage budget, not raw JSON bytes.packages/worker/src/mcp/values/service.ts#L69-L75: apply the safe serialized budget before accepting a value.packages/worker/src/mcp/values/service.node.test.ts#L377-L406: test accepted/rejected boundaries with apostrophe-heavy values.
📍 Affects 6 files
packages/shared/src/backup-restore-safety.ts#L20-L25(this comment)packages/worker/migrations/0102-restore-safe-row-sizes.sql#L12-L31packages/worker/src/app/restore-safe-row-sizes-migration.node.test.ts#L75-L109packages/worker/src/package-invocations/repo.ts#L21-L33packages/worker/src/mcp/values/service.ts#L69-L75packages/worker/src/mcp/values/service.node.test.ts#L377-L406
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/shared/src/backup-restore-safety.ts` around lines 20 - 25, Replace
the universal raw-column cap in
packages/shared/src/backup-restore-safety.ts:20-25 with a serialized SQL-row
budget that accounts for apostrophe escaping and table-specific co-resident
fields. Apply that shared budget in
packages/worker/migrations/0102-restore-safe-row-sizes.sql:12-31, including
combined text_body/html_body limits; update
packages/worker/src/package-invocations/repo.ts:21-33 and
packages/worker/src/mcp/values/service.ts:69-75 to validate escaped SQL-storage
size. Extend
packages/worker/src/app/restore-safe-row-sizes-migration.node.test.ts:75-109 and
packages/worker/src/mcp/values/service.node.test.ts:377-406 with worst-case
apostrophe-heavy and dual-large-body boundary cases.
| function boundedEmailBody(body: string | null | undefined): string | null { | ||
| if (body == null) return null | ||
| if (utf8ByteLength(body) <= maxRestorableTextColumnBytes) return body | ||
| return ( | ||
| truncateToUtf8Bytes( | ||
| body, | ||
| maxRestorableTextColumnBytes - utf8ByteLength(emailBodyTruncationNotice), | ||
| ) + emailBodyTruncationNotice |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift
Enforce a row-level budget across both email bodies.
Each column may reach 65,536 bytes, so an email with both bodies populated exceeds 131 KB before the remaining fields and SQL escaping. That can still produce an unimportable D1 INSERT. Allocate one aggregate, encoded-row-safe budget across text_body and html_body, and add coverage where both inputs are near their limits.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/worker/src/email/repo.ts` around lines 33 - 40, Replace the
independent per-column truncation in boundedEmailBody with a shared encoded-row
budget allocated across both text_body and html_body. Ensure the combined UTF-8
byte lengths, truncation notices, and both stored values remain within the
aggregate limit, while preserving null handling and valid UTF-8. Update the
email import coverage to include both bodies near their individual limits and
verify the resulting combined row stays within budget.
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@packages/worker/migrations/0105-restore-safe-row-sizes.sql`:
- Around line 21-31: Update the migration’s text_body/html_body truncation logic
to enforce a shared complete-row size budget rather than truncating each column
independently; preserve the raw MIME reference and existing marker behavior
while ensuring other row fields fit within the export limit. Add a regression
case covering an email with both bodies oversized and verify the resulting row
remains within the safe size bound.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: 3f8c181a-ef6f-450e-830d-282d19276e9f
📒 Files selected for processing (6)
packages/backup-control-plane/wrangler.jsoncpackages/worker/migrations/0105-restore-safe-row-sizes.sqlpackages/worker/src/app/restore-safe-row-sizes-migration.node.test.tspackages/worker/src/email/repo.tspackages/worker/src/mcp/values/service.node.test.tstools/migration-ledger.json
🚧 Files skipped from review as they are similar to previous changes (3)
- packages/backup-control-plane/wrangler.jsonc
- packages/worker/src/email/repo.ts
- packages/worker/src/app/restore-safe-row-sizes-migration.node.test.ts
| UPDATE email_messages | ||
| SET text_body = substr(text_body, 1, 16000) || ' | ||
| [truncated for backup-safe storage; the full message is retained in the raw MIME object]' | ||
| WHERE text_body IS NOT NULL | ||
| AND LENGTH(CAST(text_body AS BLOB)) > 65536; | ||
|
|
||
| UPDATE email_messages | ||
| SET html_body = substr(html_body, 1, 16000) || ' | ||
| [truncated for backup-safe storage; the full message is retained in the raw MIME object]' | ||
| WHERE html_body IS NOT NULL | ||
| AND LENGTH(CAST(html_body AS BLOB)) > 65536; |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift
Bound the complete email row, not each body column independently.
text_body and html_body are truncated separately to nearly 64 KiB each. If both are populated, the resulting INSERT can still exceed 100 KiB before accounting for the other columns, so the migration does not guarantee restorable exports. Apply a shared per-row budget (or clear one convenience copy when necessary) and add a regression case with both bodies oversized.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@packages/worker/migrations/0105-restore-safe-row-sizes.sql` around lines 21 -
31, Update the migration’s text_body/html_body truncation logic to enforce a
shared complete-row size budget rather than truncating each column
independently; preserve the raw MIME reference and existing marker behavior
while ensuring other row fields fit within the export limit. Add a regression
case covering an email with both bodies oversized and verify the resulting row
remains within the safe size bound.


Fixes the three root causes found while evaluating disaster recovery with live restore drills of the 2026-07-26 production export (plus the silent-failure gap):
statement too long: SQLITE_TOOBIG), and production had rows (up to 1.28 MBpackage_invocations.response_json, oldest from April 30) whose single-row INSERTs exceeded it. Every backup ever taken failed restore drills on this. A drill with statements ≤100 KB filtered out passed end-to-end (PRAGMA quick_check= ok, 66 tables), proving this was the only remaining restore blocker after fix(dr): make D1 restore imports FK-safe with foreign_keys=OFF prelude #943.packages/shared/src/backup-restore-safety.ts(64 KiB per large text column).idempotency_response_unavailableoutcome); stored email bodies are truncated (raw MIME in R2 stays canonical);value_setrejects oversized values with a storage-bucket hint.0105bounds existing rows (81 oversized invocation rows, 2 email bodies in production).<objectKey>.stats.json; oversized statements logbackup-unrestorable-statements(failure status).artifactsphase;exporter/summary.jsonwas never written, anddaily/full/is empty). The window is now 00:30–06:10 (~22 min of budget; completed days no-op via thealready-completecheck).backup-unrestorable-statements.Also refreshes
TRUSTED_RESTORE_BASELINE_SHA256(stale since #904) and documents its recipe plus the new schedule and row-size contract indocs/contributing/disaster-recovery.md.Operational note: the oversized user values have since been moved out of D1 values entirely —
solarHistoryFullnow lives as asolar_history_fulltable in the tesla-solar archive storage bucket, and the environment package caches moved to package storage — so production data is fully under the new bounds once migration0105runs.System recap — extends existing primitives (medium risk)
Mode: recap · Base:
main· Head:b8a43985Classification: extends — changes the backup control plane's finalization contract, bounds D1 row sizes at three write paths, widens the staging cron window, and adds a watchdog lane. No new primitives.
Primitives touched
backup-control-planed1-app-db0105bounds oversized rows; write-time capsemailmcp-servervalue_setrejects values over 64 KiBscheduled-crondr_export_watchdoglaneemail-blobs-r2System map
Nightly staging runs longer and gains a watchdog; the D1 backup finalization stops round-tripping through D1; write paths bound row sizes so exports import cleanly.
Legend: green = composes (wiring only) · amber = extended by this PR · red = new primitive · gray = context (unchanged, included only when an edge crosses it).
Invariants
Immutable backup objects and manifests remain append-only; the finalize change only alters how an absent manifest is verified (stored-object digest instead of a fresh D1 download). Per-user isolation is untouched — all bounded writes stay scoped to the owning
userId.Summary by CodeRabbit
New Features
Bug Fixes
Documentation