Two runtime-hardening fixes: interaction inline limit, worker readiness - #18
Conversation
… allowance Interaction artifacts (review reports, questions) must inline whole in Postgres — resume re-reads them from the artifacts row — but they were held to the store's general 64 KB threshold while the interaction schemas admit a valid report of ~248K UTF-16 units (~1.49 MB as escaped JSON; questions ~0.97 MB). A schema-valid thorough review of a large change therefore failed the run with contract_violation at the gate. The engine now passes INTERACTION_ARTIFACT_MAX_BYTES (2 MiB) as a per-call inline-limit override for artifacts that drive a checkpoint. The bound strictly dominates every schema-valid payload, so only schema-invalid or padded content can still exceed it — and that keeps failing on the existing distinct too-large contract_violation path, which fires before schema parsing. The alternative (reading interaction sources back from storage_ref) was rejected: it would promote the artifacts volume from download-convenience to run-correctness-critical and buys nothing the bound doesn't. Both previously-untested failure paths — store-time StepFailed and the resume-time legacy-row backstop — are now covered by the compliance suite, along with the >64 KB happy path and a resume leg that re-reads the report from the DB row.
…oy healthy Closes #15. worker_ok() accepted a fresh executor_registrations row, but the worker writes that BEFORE starting its pg-boss consumers — a worker that registered and then wedged inside consumer setup read as a successful deploy while nothing consumed the queue. The gap was raised P1 in two consecutive reviews of #14 and consciously deferred because closing it needs an apps/worker change. Workers now upsert a per-container row into worker_heartbeats (migration 0012; hostname inside a compose container is the container id), with consumers_ready_at written only after boss.work() has returned for every consumer, plus a 60s liveness bump from the sweeper; rows silent for a week are pruned. deploy.sh counts distinct fresh consumers-ready containers and requires one per expected WORKER_REPLICAS. The replica-count and RestartCount checks stay: each still catches what the row count cannot (out-of-band scaling; crash loops that get past consumer setup and re-freshen their row every lap). Registrations could not carry this signal — their PK is executor_id, global per executor, so with WORKER_REPLICAS > 1 one healthy replica masks the rest. The new table is the first slice of the per-worker heartbeat row deferred past M1; schema, worker write, and deploy.sh ship together so the first deploy that runs the new check has a rollback target that already writes the rows.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (16)
🚧 Files skipped from review as they are similar to previous changes (8)
📝 WalkthroughWalkthroughThe PR adds configurable inline limits for checkpoint-driving artifacts, SHA-256 patch verification, schema readiness polling, and per-container worker heartbeat readiness. Deployment verification now requires fresh consumer readiness from every expected worker replica. ChangesArtifact storage and integrity
Worker consumer readiness verification
Estimated code review effort: 4 (Complex) | ~45 minutes Sequence Diagram(s)sequenceDiagram
participant Worker
participant worker_heartbeats
participant deploy_sh
Worker->>worker_heartbeats: mark consumers ready after queue setup
Worker->>worker_heartbeats: refresh heartbeat during reconciliation
deploy_sh->>worker_heartbeats: count fresh ready containers in current fleet
worker_heartbeats-->>deploy_sh: readiness count
Possibly related PRs
🚥 Pre-merge checks | ✅ 3 | ❌ 2❌ Failed checks (2 warnings)
✅ Passed checks (3 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 2
🧹 Nitpick comments (2)
apps/worker/src/deps/readiness.ts (1)
16-29: 🗄️ Data Integrity & Integration | 🔵 Trivial | ⚡ Quick winWrite worker heartbeat timestamps from the database clock.
consumersReadyAt/heartbeatAtare currently taken from the worker process clock, while the pruning predicate and deploy readiness query compare against Postgresnow()/to_timestamp($since). A skewed worker clock can bias the deploy readiness window or prune fresh readiness rows unexpectedly. Set these fields withsqlnow that both columns allow SQL-generated values.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@apps/worker/src/deps/readiness.ts` around lines 16 - 29, Update markConsumersReady to use the database-generated current timestamp for consumersReadyAt and heartbeatAt in both the insert values and conflict-update set, using the existing SQL expression mechanism rather than the worker’s now value. Keep startedAt sourced from the process timestamp and preserve the existing pruning logic.packages/core/src/interaction-schemas.ts (1)
89-101: 📐 Maintainability & Code Quality | 🔵 TrivialKeep the 2 MiB bound tied to schema-max enforcement.
questionSchemaandreviewFindingSchemacurrently use the.max()values shown in the comment, and the arithmetic stays under 2 MiB. The warning remains useful: if any of those string maxes are changed later, update this bound/comment as well so schema-valid payloads are still dominated.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@packages/core/src/interaction-schemas.ts` around lines 89 - 101, Keep INTERACTION_ARTIFACT_MAX_BYTES synchronized with the .max() limits enforced by questionSchema and reviewFindingSchema. If those schema maximums change, recalculate and update the 2 MiB bound and its explanatory comment so every schema-valid payload remains within the limit.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@docs/design/01-domain-model.md`:
- Line 225: Clarify the artifact storage boundary: in
docs/design/01-domain-model.md (225-225), state that the 64 KiB inline rule
applies only to small text/JSON artifacts, while path-backed kind === "file"
artifacts use storageRef even below that size. Apply the corresponding
small-file artifact volume wording and backup-survival qualification in
docs/manual/en/06-operations.md (82-82, 129-129) and mirror both changes in
docs/manual/zh-CN/06-operations.md (81-81, 128-128), keeping the English and
zh-CN manuals synchronized.
In `@infra/deploy.sh`:
- Around line 362-367: Update the worker readiness query in the HEALTH_TIMEOUT
verification loop so it does not require consumers_ready_at to be newer than the
current deploy’s $since timestamp. Use boot-relative readiness by requiring
consumers_ready_at to correspond to the worker’s current started_at and
confirming heartbeat_at is fresh, while preserving the existing expected-count
and missing-table handling.
---
Nitpick comments:
In `@apps/worker/src/deps/readiness.ts`:
- Around line 16-29: Update markConsumersReady to use the database-generated
current timestamp for consumersReadyAt and heartbeatAt in both the insert values
and conflict-update set, using the existing SQL expression mechanism rather than
the worker’s now value. Keep startedAt sourced from the process timestamp and
preserve the existing pruning logic.
In `@packages/core/src/interaction-schemas.ts`:
- Around line 89-101: Keep INTERACTION_ARTIFACT_MAX_BYTES synchronized with the
.max() limits enforced by questionSchema and reviewFindingSchema. If those
schema maximums change, recalculate and update the 2 MiB bound and its
explanatory comment so every schema-valid payload remains within the limit.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: a460db5a-2655-44f6-903f-60e699b335c3
📒 Files selected for processing (23)
CHANGELOG.mdapps/worker/src/deps/artifacts.test.tsapps/worker/src/deps/artifacts.tsapps/worker/src/deps/readiness.test.tsapps/worker/src/deps/readiness.tsapps/worker/src/index.tsdocs/design/01-domain-model.mddocs/design/04-execution-runtime.mddocs/design/08-deployment.mddocs/manual/en/06-operations.mddocs/manual/zh-CN/06-operations.mdinfra/deploy.shinfra/deploy.test.tspackages/core/src/interaction-schemas.tspackages/db/drizzle/0012_worker_heartbeats.sqlpackages/db/drizzle/meta/0012_snapshot.jsonpackages/db/drizzle/meta/_journal.jsonpackages/db/src/schema/registry.tspackages/db/src/schema/runs.tspackages/orchestration/src/engine/deps.tspackages/orchestration/src/engine/engine.integration.test.tspackages/orchestration/src/engine/engine.tspackages/orchestration/src/engine/fakes.ts
…ailing publish The git.push evidence check compared the fresh workspace snapshot against artifactValues, which holds "" for any patch artifact past the 64 KB inline threshold — so every run whose reviewed diff exceeded 64 KB died at publish with a phantom "workspace changed after the reviewed evidence". This was the remaining follow-up from PR #5's review rounds, and the sibling of the interaction-artifact fix earlier on this branch: same threshold, different consumer. A patch cannot get the interaction artifacts' raised inline allowance — patches are capped at 25 MB, far past anything Postgres should inline — so the check instead reads the stored bytes back via a new ArtifactStore.read(storageRef). The stored patch IS the approved evidence: drifted workspaces still fail exactly as before, and evidence that cannot be read back (lost volume, corrupted row) fails the push with a distinct contract_violation rather than publishing unverified. DiskArtifactStore.read refuses refs outside the storage root, so a corrupted row cannot become an arbitrary-file-read primitive; the in-memory fake persists spilled content across engine legs the way the real artifacts volume does.
…ness, replicas guard Four findings from the codex review of this branch, all confirmed: P1: patch evidence read back from the artifacts volume was mutable after approval — the volume is agent-writable (same container, same bun user, no functional inner sandbox under Docker), so an agent-left process could rewrite workspace and spilled patch together and publish unapproved changes. Every stored artifact row now carries a store-time sha256 (migration 0013) and git.push verifies a larger-than-inline patch against that digest, which lives in Postgres where agent subprocesses cannot reach; ArtifactStore.read() is removed and disk bytes are never trusted as evidence. A spilled patch without a digest fails as unverifiable. P1: rollbacks and same-SHA redeploys that reuse unchanged worker containers never produce a fresh consumers_ready_at (the stamp is boot-only, unlike the sweeper-refreshed registration it replaced), so worker_ok() falsely reported the instance down. The predicate is now ready AND alive: consumers_ready_at non-null plus a fresh sweeper heartbeat — reused healthy containers converge within one 60s tick. P1: a stale readiness stamp could survive a restart and mask a boot that wedges in consumer setup. Every boot now writes a boot-start transition first (markBootStarted: startedAt reset, consumersReadyAt cleared), so readiness always describes the current boot. P2: WORKER_REPLICAS=0 passed verification vacuously (0==0, 0>=0); deploy.sh now rejects values below 1 before building anything.
…, written-bytes digest Four findings from the second codex review of this branch. Three are fixed; the first is confirmed and documented as the accepted residual it restates. P1 (accepted, documented): the evidence digest is reachable from a read-write agent that recovers worker credentials via /proc/1/environ — same UID, same PID namespace, no functional inner sandbox under Docker, and Yama does not guard PTRACE_MODE_READ. Confirmed. No in-place hardening can close it: with DATABASE_URL an attacker forges approvals outright (the worker role must keep UPDATE on checkpoints for expiry), and overrides digests without UPDATE because resume replays artifact rows by createdAt. This is the documented container-is-the-boundary posture (design 08, Top Risks #2); comments, design docs, and the CHANGELOG now say "tamper-resistance within the posture, not a boundary" instead of implying agents cannot reach Postgres. Per-run isolation stays the M2 work (issue draft prepared). P1: worker_ok() counted any ready row with a fresh heartbeat, so a foreign container on the same database — debug docker run, out-of-band scale leftover, second stack, or an old container beating just before replacement — could satisfy the count while a replacement wedged. The query is now scoped to the hostnames of this fleet's running containers ({{.Config.Hostname}}, by construction not convention), and aliveness uses a sliding 90s window: the fixed post-deploy stamp let a single beat mask a later wedge for the rest of verification. The sweeper now beats first in its tick so a failing sweep cannot skip the heartbeat. P1: the worker's boot-time insert into worker_heartbeats raced the api's on-boot migration on exactly the deploy that ships the table; the resulting crash-loop moved RestartCount and rolled back a good deploy. The compose worker now gates on api health (depends_on service_healthy — the api is healthy only after migrate/seed/publish complete; mirrors the VM unit's /healthz ExecStartPre), and both `up -d` call sites are bounded by timeout 600, because the gate makes compose's wait open-ended when an api crash-loops (start_period resets per restart) and an unbounded hang would run to the Janus unit's SIGKILL with no rollback. P2: file-kind artifact digests were computed by re-reading the mutable source after the copy; the spill path now hashes and size-counts the exact byte stream being written in a single pass, enforcing the size cap mid-stream and deleting partial files on abort.
…heartbeats CodeRabbit's review of the first pushed state raised two actionable comments and two nitpicks; its other Major (boot-only readiness fails same-commit redeploys) was already fixed by the previous two commits, which replaced the post-deploy ready-stamp requirement with the ready-and-alive predicate it suggests. - The "≤64 KB inline" wording in the domain model and both manuals overclaimed: the store inlines only text/JSON artifacts, so a file-kind artifact lives on the artifacts volume at ANY size. An operator planning backups from those sentences would have believed small file artifacts survive volume loss. All five call-out sites now state the file-kind boundary, both locales together. - worker_heartbeats timestamps are now written with the database clock (now()) instead of the worker process clock: the deploy verification window and the prune predicate compare against Postgres now(), and a skewed worker clock must not shift rows in or out of either. - The INTERACTION_ARTIFACT_MAX_BYTES derivation comment now says explicitly that changing any schema .max() requires re-deriving it. - Test-infra hardening found while verifying: every suite's "drop schema public cascade" is now IF EXISTS-guarded — a run that dies between drop and create (seen locally via transient macOS setsockopt failures killing new connections) previously left the test DB without a public schema, cascading failures into every later run.
There was a problem hiding this comment.
Actionable comments posted: 3
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
apps/worker/src/deps/readiness.ts (1)
36-49: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick winDecouple the stale-row prune from the essential readiness write.
markConsumersReadyperforms two unrelated operations sequentially: the readiness upsert (essential) and a housekeeping DELETE (best-effort). TheINSERT ... ON CONFLICTstatement commits immediately, so a failure in the DELETE that follows cannot undo the readiness write. But because both statements run under oneawaitchain without atry/catch, a DELETE failure (lock contention, statement timeout) makes the whole function reject.At the call site (
apps/worker/src/index.ts, line 190 perawait markConsumersReady(db, containerId);), this rejection propagates afterboss.work()has already started every consumer. If nothing catches it, the worker process can crash on a working set of consumers, andinfra/deploy.sh'sworker_ok()RestartCount check would then read that as crash-looping and could roll back an otherwise healthy deploy.Wrap the prune in its own
try/catchso a housekeeping failure cannot fail the readiness signal that consumer setup already earned.🛠️ Proposed fix
export async function markConsumersReady(db: Db, containerId: string): Promise<void> { await db .insert(workerHeartbeats) .values({ containerId, startedAt: DB_NOW, consumersReadyAt: DB_NOW, heartbeatAt: DB_NOW }) .onConflictDoUpdate({ target: workerHeartbeats.containerId, set: { consumersReadyAt: DB_NOW, heartbeatAt: DB_NOW }, }); // containers are recreated on every deploy, so rows accumulate one per // container forever; a week of silence is far past any freshness window - await db - .delete(workerHeartbeats) - .where(lt(workerHeartbeats.heartbeatAt, sql`now() - interval '7 days'`)); + try { + await db + .delete(workerHeartbeats) + .where(lt(workerHeartbeats.heartbeatAt, sql`now() - interval '7 days'`)); + } catch (err) { + // readiness above already committed; a prune failure must not fail startup + console.error("markConsumersReady: stale-row prune failed", err); + } }🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@apps/worker/src/deps/readiness.ts` around lines 36 - 49, Update markConsumersReady so the stale workerHeartbeats DELETE runs inside its own try/catch after the readiness upsert; preserve the upsert’s error propagation while swallowing or reporting prune failures without allowing them to reject the function.
🧹 Nitpick comments (1)
apps/worker/src/deps/artifacts.test.ts (1)
171-201: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winAdd coverage for the failed-write cleanup path.
These tests cover the success paths. The new failure path in
streamToDiskis untested: a file that exceedsAGRIPPA_MAX_ARTIFACT_BYTESduring the copy must throwArtifactTooLargeErrorand must leave no partial file at the storage reference. SetAGRIPPA_MAX_ARTIFACT_BYTESbelow the file size, then assert both the thrown error and the absence of the partial output.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@apps/worker/src/deps/artifacts.test.ts` around lines 171 - 201, Add a test covering the oversized file failure path in streamToDisk: set AGRIPPA_MAX_ARTIFACT_BYTES below the source file size, assert store.store throws ArtifactTooLargeError, and verify no partial file remains at the returned or expected storage reference. Keep the existing success-path assertions unchanged.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@apps/worker/src/deps/artifacts.ts`:
- Around line 102-127: Update streamToDisk so each writer.write(chunk) result is
captured and awaited when it returns a pending Promise before processing the
next chunk. Preserve the existing size validation, hashing, cleanup behavior,
and final writer.end() call while ensuring buffered FileSink writes are drained
before advancing.
In `@docs/manual/en/06-operations.md`:
- Line 82: Update the ARTIFACT_STORAGE_ROOT entry in the operations
documentation to use the precise 64 KiB threshold and clarify that all file-kind
artifacts are stored there while remaining subject to the
AGRIPPA_MAX_ARTIFACT_BYTES per-artifact size cap.
- Line 129: Update the artifact-store loss description in the operations
documentation to remove the claim that publish-time patch verification is
unaffected. Distinguish that Postgres retains digests and metadata, while
verification still requires readable spilled evidence and must fail when that
evidence is unavailable.
---
Outside diff comments:
In `@apps/worker/src/deps/readiness.ts`:
- Around line 36-49: Update markConsumersReady so the stale workerHeartbeats
DELETE runs inside its own try/catch after the readiness upsert; preserve the
upsert’s error propagation while swallowing or reporting prune failures without
allowing them to reject the function.
---
Nitpick comments:
In `@apps/worker/src/deps/artifacts.test.ts`:
- Around line 171-201: Add a test covering the oversized file failure path in
streamToDisk: set AGRIPPA_MAX_ARTIFACT_BYTES below the source file size, assert
store.store throws ArtifactTooLargeError, and verify no partial file remains at
the returned or expected storage reference. Keep the existing success-path
assertions unchanged.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: 8aacc5f8-d2e6-4942-8f89-e5983674a089
📒 Files selected for processing (26)
CHANGELOG.mdapps/api/src/test/helpers.tsapps/worker/src/deps/artifacts.test.tsapps/worker/src/deps/artifacts.tsapps/worker/src/deps/readiness.test.tsapps/worker/src/deps/readiness.tsapps/worker/src/deps/workspace.test.tsapps/worker/src/index.tsdocs/design/01-domain-model.mddocs/design/04-execution-runtime.mddocs/design/08-deployment.mddocs/manual/en/06-operations.mddocs/manual/zh-CN/06-operations.mdinfra/deploy.shinfra/deploy.test.tsinfra/docker-compose.ymlinfra/env/.env.examplepackages/core/src/interaction-schemas.tspackages/db/drizzle/0013_artifact_sha256.sqlpackages/db/drizzle/meta/0013_snapshot.jsonpackages/db/drizzle/meta/_journal.jsonpackages/db/src/schema/runs.tspackages/orchestration/src/engine/deps.tspackages/orchestration/src/engine/engine.integration.test.tspackages/orchestration/src/engine/engine.tspackages/orchestration/src/engine/fakes.ts
🚧 Files skipped from review as they are similar to previous changes (6)
- docs/design/01-domain-model.md
- packages/core/src/interaction-schemas.ts
- packages/db/src/schema/runs.ts
- docs/manual/zh-CN/06-operations.md
- packages/orchestration/src/engine/fakes.ts
- CHANGELOG.md
|
|
||
| 1. The **database** — Compose: the `pgdata` volume; VM: `pg_dump agrippa` — schedule per your policy. | ||
| 2. The **artifact store** — Compose: the `artifacts` volume; VM: `/var/lib/agrippa/artifacts`. Losing it loses downloads over 64 KB (metadata and small artifacts survive in Postgres). | ||
| 2. The **artifact store** — Compose: the `artifacts` volume; VM: `/var/lib/agrippa/artifacts`. Losing it loses downloads of text artifacts over 64 KB and of `file`-kind artifacts of any size (metadata, small text artifacts, and checkpoint-driving artifacts survive in Postgres; publish-time patch verification uses digests stored in Postgres, so it is unaffected). |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win
Do not claim patch verification survives artifact-store loss.
The digest and metadata survive in Postgres, but verification still needs to read spilled evidence. The PR contract states that unreadable evidence fails verification. Replace “so it is unaffected” with wording that distinguishes digest retention from evidence availability.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@docs/manual/en/06-operations.md` at line 129, Update the artifact-store loss
description in the operations documentation to remove the claim that
publish-time patch verification is unaffected. Distinguish that Postgres retains
digests and metadata, while verification still requires readable spilled
evidence and must fail when that evidence is unavailable.
…h gate Third codex review, one blocking finding: deploy.sh edits take effect one deploy LATER (the script runs from the previous deploy's inode; documented at its own line 28), while compose reads the freshly reset tree. So the deploy that ships this PR would get the new `depends_on: api: service_healthy` gate with the OLD, unbounded `up -d` — and this PR ships two migrations, the most likely trigger: a failing migration exits the api, `restart: unless-stopped` relaunches it, health resets to `starting` under a fresh 180s start_period every lap, and compose's dependency wait (which only errors on exited/unhealthy) effectively never terminates. The Janus unit then SIGKILLs at TimeoutStartSec with no rollback and the tree left on the failed commit. Fixed by moving the ordering into the worker image, which IS rebuilt and started by the old script on this very deploy: awaitSchema() compares the image's migration journal against drizzle.__drizzle_migrations and waits before any DB write. That is strictly broader than the compose gate — it also covers host reboots and `docker start` (restart policies ignore depends_on), the VM topology, and AGRIPPA_MIGRATE_ON_BOOT=0, where api health says nothing about the schema — and `up -d` no longer blocks on another service's health at all. The journal comparison is used rather than probing for a named table because a column-only migration (0013) would sail straight past a table-existence check. Its 300s bound deliberately exceeds deploy.sh's HEALTH_TIMEOUT so a schema that never arrives fails the deploy as "worker never became ready" instead of a restart-count mismatch that reads like a crash loop; a cross-file test pins that ordering. The `up -d` timeouts drop to 300s, sized so the worst path (up + verify + rollback's up + verify) fits TimeoutStartSec=1800 after a slow build. Also from the same review round: - deps.ts still claimed agents "cannot reach" Postgres — the one over-claim missed last round; now posture-level tamper resistance, not a boundary. - Bun.FileSink.write() returns a Promise under backpressure; streamToDisk ignored it, so end() could run before those chunks landed (CodeRabbit). - Docs said 64 KB where the code means 64 KiB, and "file artifacts of any size" read as unbounded next to AGRIPPA_MAX_ARTIFACT_BYTES (CodeRabbit). - worker_ok() sanitized container hostnames by STRIPPING unexpected characters, so a hostname carrying - or . would silently stop matching what os.hostname() wrote and roll back a healthy deploy; it now accepts that charset and fails closed on anything else. CodeRabbit's remaining Major — that the backup note must not say publish verification is unaffected by artifact-store loss — is a stale premise: since the digest change, verification hashes the current workspace diff and compares it to the Postgres digest, never reading the volume. The note keeps its claim and now states that mechanism.
Three runtime fixes, one commit each. The first and third are siblings — both halves of the 64 KB inline threshold killing legitimate runs — and the second closes #15.
1.
fix(engine): checkpoint-driving artifacts get their own 2 MiB inline allowanceInteraction artifacts (review reports, questions) must inline whole in Postgres — resume re-reads them from the artifacts row — but they were held to the store's general 64 KB threshold while the interaction schemas admit a valid report of ~248K UTF-16 units (~1.49 MB as escaped JSON). A schema-valid thorough review of a large change therefore killed its run with
contract_violationat the review gate.The engine now passes
INTERACTION_ARTIFACT_MAX_BYTES(2 MiB) as a per-call inline-limit override for artifacts that drive a checkpoint. The bound provably dominates every schema-valid payload (derivation in the constant's comment), so only schema-invalid or padded content can still exceed it — and that keeps failing on the existing distinct too-largecontract_violation, which fires before schema parsing.New compliance-suite coverage: the >64 KB happy path, the beyond-limit failure, a resume leg that re-reads the report from the DB row, and the previously-untested legacy-row backstop (
inline=null+storage_refrows predating store-time validation).2.
fix(deploy): prove each worker replica consumes before calling a deploy healthyCloses #15.
worker_ok()accepted a freshexecutor_registrationsrow, which the worker writes before starting its pg-boss consumers — a worker that registered and then wedged inside consumer setup read as a successful deploy while nothing consumed the queue.Workers now upsert a per-container row into the new
worker_heartbeatstable (migration0012; hostname inside a compose container is the container id), withconsumers_ready_atwritten only afterboss.work()has returned for every consumer, plus a 60 s liveness bump from the sweeper (rows silent for a week are pruned).deploy.shcounts distinct fresh consumers-ready containers and requires one per expectedWORKER_REPLICAS. The replica-count and RestartCount checks stay — each catches what the row count cannot (out-of-band scaling; crash loops that get past consumer setup). Registrations could not carry the signal: their PK isexecutor_id, global per executor, so one healthy replica masks the rest. The table is the first slice of the per-worker heartbeat row deferred past M1.Schema, worker write, and
deploy.shship in one commit deliberately: deploy-script changes take effect one deploy later, so the first deploy that runs the new check has a rollback target that already writes the rows.3.
fix(engine): read big patch evidence back from the store instead of failing publishThe remaining follow-up from #5's review rounds. The
git.pushevidence check compared the fresh workspace snapshot againstartifactValues, which holds""for any patch artifact past the 64 KB inline threshold — so every run whose reviewed diff exceeded 64 KB died at publish with a phantom "workspace changed after the reviewed evidence".A patch cannot get fix 1's raised inline allowance (patches are capped at 25 MB), so the check instead reads the stored bytes back via a new
ArtifactStore.read(storageRef)— the stored patch is the approved evidence. Drifted workspaces still fail exactly as before; evidence that cannot be read back (lost volume, corrupted row) fails the push with a distinctcontract_violationrather than publishing unverified;DiskArtifactStore.readrefuses refs outside the storage root so a corrupted row cannot become an arbitrary-file-read primitive. Compliance coverage: big-patch publish via read-back, big-patch drift, unreadable evidence.Verification
Full gate green at each commit:
bun run check,bun test(local Postgres up — 336 pass / 0 fail, integration suites ran),bun run templates:validate,bun run build, plusbun run db:migratefor0012, andbunx commitlint --from mainover the range. Real-host proof for the deploy check lands on the second deploy after merge (script-takes-effect-next-deploy).Summary by CodeRabbit
New Features
Bug Fixes