Skip to content

fix(cache,query,jobs,core,db)!: a cached query answered the wrong tenant, and five ways an interrupted process never recovered - #110

Merged
sebyx07 merged 2 commits into
mainfrom
fix/cache-query-jobs-lifecycle
Aug 17, 2026
Merged

sebyx07 merged 2 commits into
mainfrom
fix/cache-query-jobs-lifecycle

Conversation

@sebyx07

@sebyx07 sebyx07 commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Slices 02 (tier 2–3 bugs) and 06 (concurrency & lifecycle) from
docs/plans/2026/08/16/101-deep-dive-bug-audit —
the non-realtime half. The realtime half landed as #107.

Four agents on disjoint package sets in one checkout: packages/jobs · packages/cache ·
packages/query + packages/cli · packages/core + packages/db.

The one that matters

A cache: query served one actor's rows to the next. cacheKeyFor returned
query:<name>:<fingerprint(input)>:<tags> — no actor, no tenant — while sql(input, ctx) is handed
the Ctx and @ultimat3/entity derives every tenant predicate from ctx.actor.orgId rather than
from the input. Reproduced: a query declaring cache: { tags: [], ttlMs: 60_000 } and filtering on
ctx.actor.orgId answered an org-b actor with {id:'a1', orgId:'org-a', secret:'ALPHA'}.
Cross-tenant disclosure with no attacker: two logged-in users and one cached query.

The key now carries the read's authority — query:<name>:<authority>:<fingerprint>:<tags>, the
authority being JSON.stringify([kind, id, orgId ?? null]). JSON rather than a joined string
because an actor id is app data that may contain the separator, and a value that can spell a
boundary can spell someone else's — the rule @ultimat3/entity's scopeKey already states. A new
cache.scope picks the sharing width: 'actor' (default), 'tenant', 'global'. The default
is the narrowest, which is what makes forgetting safe; a fourth scope is an assertNever compile
error. The request memo was never affected — it keys on Ctx identity.

Findings, one row each

# Package Defect Evidence
1 query cache key omitted the actor → cross-tenant read reproduced, cache-authority.test.ts
2 cli/query read cache installed beside the tier registry, so cache.invalidates busted nothing without Redis; invalidateQueryTags had zero production callers dev-cache.test.ts
3 cache/query an invalidation landing mid-read was overwritten by pre-write rows for the full TTL invalidation-race.test.ts, cache-fence.test.ts
4 cache Redis bust dropped the tag bucket atomically with SMEMBERS, orphaning members on a refused DEL live Redis
5 cache set wrote the value before the tag SADD — and reversing it alone is worse live Redis, redis-ordering.test.ts
6 cache invalidation fanned out in read order, so a racing read promoted a stale value backwards tiers.test.ts
7 jobs a rejecting fleetSlots.acquire() leaked the limiter lease → the worker role died permanently and silently worker-fleet-slots.test.ts
8 jobs the fleet heartbeat never acted on a lost slot; the comment claimed it did new X_JOB_SLOT_LOST
9 jobs options.context() ran after the heartbeat and slot timers started — app code that throws left both running forever worker-run.ts
10 jobs OutboxRelay.stop() returned underneath the pass in flight outbox.test.ts
11 jobs x jobs ls against x dev paged the oldest hundred rows — the limit landed after the sort driver-parity.test.ts (new mechanism)
12 core the drain deadline was opt-in, and the two roles that most need it (jobs, realtime) declared none lifecycle.test.ts
13 db a pglite observer straggler ran against a closed handle after the transaction settled pglite-observer.test.ts
14 query fingerprint was FNV-1a/32 over client-chosen input, backing both cursors and the shared cache key stable.test.ts
15 cache semantic.remember bypassed assertTtl semantic.test.ts
16 cli nothing in the repo caught a dropped await noFloatingPromises enforced
17 cli the deployed app was linted against nothing — dummy/social-media-clone/biome.json set "root": false with no extends proven with a planted probe
18 scripts four repo-scanning gate tests timed out under the gate's own 8-way sharding reproduced under --workers 8, twice

The intermittent, finally explained

An unexplained shard failure recurred five times across this audit. It is four scripts/ tests
paying a whole-repo cost against bun's default 5000ms budget while eight shards compete for the
same cores — and which shard a file lands in depends on the file count, so it read as flake rather
than as a slow test. It surfaced now because this PR adds files: every repo-scanning test got
slower.

The repo had already diagnosed this once and fixed it for error-contract.test.ts, whose comment
ends "Same shape as scripts/verify.test.ts" — naming one of the files that kept failing. The
diagnosis was written down and never applied. Both collectSourceFiles tests moved, not only the
one observed failing, because they call the identical scan and fixing one relocates the failure.
Proven by reproducing under the gate's own command and re-running after the fix, once plain and
once with four extra CPU hogs: 16 of 16 shard processes exit 0.

One correction it forced on my own brief: verify.test.ts's cost is not a directory walk. It runs
on a small temp dir; the expense is registeredErrorCodes() dynamically importing all 29 packages
regardless.

Premises the hive falsified

Nine brief premises were wrong and the agents said so with the line. The ones that changed the fix:

  • "Reverse the Redis set ordering." Half a fix, and worse alone: SADD-then-SET with no
    re-check lets the bust SREM the membership and the later SET publish a row unreachable by
    any tag — permanently uninvalidatable. Shipped with an SISMEMBER re-check that deletes the
    value it just wrote when a bucket says it is gone, and only a literal 0 counts as evidence.
  • "Fix queryHash." Too narrow. fingerprint also backs cacheKeyFor, so fixing only the
    cursor path would have left a 32-bit hash over client-chosen input as the shared cache entry
    key — strictly worse than the case named.
  • "Repeated suspensions drive attempt negative." Impossible: nack is fenced on
    state === 'running', which only claim sets, and claim increments.
  • "configureLifecycle({deadlineMs}) is effectively unused." False — http/src/server.ts:97
    declares one on every createServer.
  • "X_SHUTDOWN_TIMEOUT is new." Already shipped at core/src/error-codes.ts:61.
  • "bestEffort is exported from tier-failures.ts." It was not exported at all, and its
    TierName parameter has no rung for query's read tier — hence TierLabel.
  • "Bare Date.now() in four cache files." The list is empty; fixed in an earlier PR.

One agent also overturned my call: I leaned "loud is usually right" on the pglite straggler and D
showed with evidence that it is not. And D caught its own overclaim — it reported the whole-drain
budget as proved when the mutation passed 18/18; it was implemented but unenforced until the test at
lifecycle.test.ts:263.

The coordinator decision

The drain deadline is now bounded by default (25s) rather than opt-in. D built it opt-in as
briefed, then showed that jobs and realtime declare no budget — so the proven symptom, a worker
pod SIGKILLed mid-job, stayed unfixed. Shipping a mechanism that does not fix the symptom while
claiming the deadline works is a claim wider than what is enforced: the third instance of that class
this run, after "time-to-consistent" and "0 lost". The flip broke nothing — jobs 411 pass, realtime
675 pass.

Read X_SHUTDOWN_TIMEOUT literally: a hook past the deadline is abandoned, not stopped. It is
still running when the process exits. The framework cannot cancel app code it did not write, and
both fix: lines now name the pair that must move together — configureLifecycle({ deadlineMs })
and terminationGracePeriodSeconds.

Enforcement gap closed in-PR

Nothing caught a dropped await. Measured noFloatingPromises at 2 violations across 3560 files in
7s and enabled it rather than deferring (bun run lint 6.0s → 7.1s). Enabling it is what exposed
finding 17: the deployed demo app inherited none of the repo's lint rules.

Breaking

Six, all in CHANGELOG.md under ### Changed. [Unreleased] already carries a dozen BREAKING
entries from #105–#107, so the next release is a major regardless and none of these forces a new
decision.

drainDeadlineMs() returns number always · cacheKeyFor takes a required fourth argument ·
fingerprint is SHA-256/16 so pre-existing cursors are rejected once as X_CURSOR_INVALID ·
semantic.remember rejects a TTL the tiers would reject · OutboxRelay.stop() returns
Promise<void> · TierFailure.tier widens to TierLabel.

Codes

X_JOB_SLOT_LOST, X_QUERY_CACHE_TTL_INVALID — both in wiki/Error-Codes.md, manifest
regenerated (389 codes, 29 packages).

Known gaps recorded, not fixed

  • The cache fill fence is per process — two pods can still interleave a load on one with a write
    and bust on the other. Cross-node needs a Redis-side epoch and a wire change.
  • createCacheStack has zero production callers, so its copy of the fence is dormant. The live
    fenced path is runQuery → readRows → readThrough → fill.
  • recentTierFailures() has no reader outside cache, and /_x's invalidations source is
    unwired — routed to the tiers 4–5 slice.
  • invalidateWireTags has zero callers repo-wide, but it is not dead: it is the wire-form door
    x cache bust will call, and receiveInvalidationBroadcast is the sibling door that already has
    a production caller. Deleting it would leave the planned command with no entry point. Recorded
    for the dead-code slice to weigh, not fixed here.

Gate

bun run verify — 17 steps. bun run scripts/reference-app-gate.ts — both tracked apps on their
ratchet.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

sebyx07 and others added 2 commits August 16, 2026 23:51
…ant, and five ways an interrupted process never recovered

Slices 02 and 06 of the deep-dive bug audit, minus the realtime half that
landed as #107. Four agents on disjoint package sets in one checkout.

The one that matters: `cacheKeyFor` keyed on name + input + tags and never on
the caller, while `sql(input, ctx)` derives every tenant predicate from
`ctx.actor.orgId`. Reproduced — an `org-b` actor was answered with
`{id:'a1', orgId:'org-a', secret:'ALPHA'}`. The key now carries the read's
authority, and a new `cache.scope` defaults to `'actor'`: the narrowest, which
is what makes forgetting it safe.

Three defects had to land together or make things worse. Reversing the Redis
`set` ordering alone lets a bust `SREM` the membership while the later `SET`
publishes a row unreachable by any tag; it ships with an `SISMEMBER` re-check.
Fixing the cursor hash alone would have left a 32-bit hash over client-chosen
input as the shared cache key. And an invalidation racing a read was invisible
until the read cache lived inside the tier registry rather than beside it.

The drain deadline flipped from opt-in to bounded by default at 25s. It was
built opt-in as briefed, but `jobs` and `realtime` declare no budget, so the
proven symptom — a worker pod SIGKILLed mid-job — stayed unfixed. Read
X_SHUTDOWN_TIMEOUT literally: a hook past the deadline is abandoned, not
stopped, and both fix: lines now name `terminationGracePeriodSeconds` beside
`configureLifecycle`.

Enforcement gap closed in passing: nothing caught a dropped `await`.
Enabling `noFloatingPromises` cost 1.1s of lint and exposed that
`dummy/social-media-clone` — the deployed app — set `"root": false` with no
`extends`, so it was linted against nothing at all.

BREAKING CHANGE: `drainDeadlineMs()` returns `number` always; `cacheKeyFor`
takes a required fourth `authority` argument; `fingerprint` is SHA-256/16, so
cursors minted before this are rejected once as X_CURSOR_INVALID;
`semantic.remember` rejects a TTL the tiers would reject; `OutboxRelay.stop()`
returns `Promise<void>`; `TierFailure.tier` widens to `TierLabel`.

New codes: X_JOB_SLOT_LOST, X_QUERY_CACHE_TTL_INVALID.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RBwWKBJkiogA4mDaJiJf3D
…o saying 3 skipped

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RBwWKBJkiogA4mDaJiJf3D
@coderabbitai

coderabbitai Bot commented Aug 17, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 57 minutes

Limit details: You’ve used all 1 included review currently available under your plan. You completed 78 included PR reviews in the past 7 days; at that activity level, included reviews refill at 1 review per hour.

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yml

Review profile: ASSERTIVE

Plan: Pro

Run ID: f9030b10-6956-4bfd-8779-e5c96fd95ecd

📥 Commits

Reviewing files that changed from the base of the PR and between bde3cd3 and e11ccc6.

📒 Files selected for processing (82)
  • CHANGELOG.md
  • biome.json
  • docs/plans/2026/08/16/101-deep-dive-bug-audit/status.yml
  • dummy/social-media-clone/biome.json
  • framework.manifest.json
  • packages/cache/CLAUDE.md
  • packages/cache/README.md
  • packages/cache/src/fence.test.ts
  • packages/cache/src/fence.ts
  • packages/cache/src/index.ts
  • packages/cache/src/invalidate.ts
  • packages/cache/src/invalidation-race.test.ts
  • packages/cache/src/redis-fake.ts
  • packages/cache/src/redis-ordering.test.ts
  • packages/cache/src/redis.live.test.ts
  • packages/cache/src/redis.test.ts
  • packages/cache/src/redis.ts
  • packages/cache/src/semantic.test.ts
  • packages/cache/src/semantic.ts
  • packages/cache/src/set-options.ts
  • packages/cache/src/single-flight.test.ts
  • packages/cache/src/single-flight.ts
  • packages/cache/src/tier-failures.test.ts
  • packages/cache/src/tier-failures.ts
  • packages/cache/src/tiers.test.ts
  • packages/cache/src/tiers.ts
  • packages/cli/src/dev-cache.test.ts
  • packages/cli/src/dev-cache.ts
  • packages/cli/src/dev-roles-csp.test.ts
  • packages/cli/src/dev-roles-fixture.ts
  • packages/cli/src/dev-roles-identity.test.ts
  • packages/cli/src/dev-roles.test.ts
  • packages/cli/src/dev-roles.ts
  • packages/core/CLAUDE.md
  • packages/core/src/index.ts
  • packages/core/src/lifecycle-deadline.ts
  • packages/core/src/lifecycle.test.ts
  • packages/core/src/lifecycle.ts
  • packages/db/CLAUDE.md
  • packages/db/src/fake-pglite.ts
  • packages/db/src/pglite-observer.test.ts
  • packages/db/src/pglite.test.ts
  • packages/db/src/pglite.ts
  • packages/db/src/transaction.ts
  • packages/jobs/CLAUDE.md
  • packages/jobs/README.md
  • packages/jobs/src/driver-memory.ts
  • packages/jobs/src/driver-parity.test.ts
  • packages/jobs/src/errors.ts
  • packages/jobs/src/index.ts
  • packages/jobs/src/outbox.test.ts
  • packages/jobs/src/outbox.ts
  • packages/jobs/src/run-signal.test.ts
  • packages/jobs/src/run-signal.ts
  • packages/jobs/src/task.test.ts
  • packages/jobs/src/worker-fleet-slots.test.ts
  • packages/jobs/src/worker-fleet-slots.ts
  • packages/jobs/src/worker-run.ts
  • packages/jobs/src/worker.ts
  • packages/query/CLAUDE.md
  • packages/query/README.md
  • packages/query/src/cache-authority.test.ts
  • packages/query/src/cache-degraded.test.ts
  • packages/query/src/cache-fence.test.ts
  • packages/query/src/cache.test.ts
  • packages/query/src/cache.ts
  • packages/query/src/errors.ts
  • packages/query/src/index.ts
  • packages/query/src/live.test.ts
  • packages/query/src/query.test.ts
  • packages/query/src/query.ts
  • packages/query/src/read-cache.test.ts
  • packages/query/src/read-cache.ts
  • packages/query/src/read.test.ts
  • packages/query/src/read.ts
  • packages/query/src/stable.test.ts
  • packages/query/src/stable.ts
  • scripts/boundaries.test.ts
  • scripts/manifest.test.ts
  • scripts/verify.test.ts
  • wiki/Error-Codes.md
  • wiki/Known-Gaps.md
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/cache-query-jobs-lifecycle

Comment @coderabbitai help to get the list of available commands.

@sebyx07
sebyx07 merged commit d44fa95 into main Aug 17, 2026
5 checks passed
@sebyx07
sebyx07 deleted the fix/cache-query-jobs-lifecycle branch August 17, 2026 04:55
sebyx07 added a commit that referenced this pull request Aug 17, 2026
auth and entity needed nothing — every finding naming them was already closed by
#104 and #106, verified file-by-file before #112 was scoped. That is what took
the PR from an estimated ~70 files to 43.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RBwWKBJkiogA4mDaJiJf3D
sebyx07 added a commit that referenced this pull request Aug 17, 2026
…her, and ten codes paged the on-call for a caller's mistake (#112)

* fix(action,http,core)!: one caller's idempotent response went to another, and ten codes paged the on-call for a caller's mistake

The http and action remainder of audit slices 02 and 06, plus the gate step
slice 02 asks for by name. auth and entity are absent because they are already
closed — every finding naming them landed in #104 and #106, verified before
this wave was scoped.

`idempotencyKeyFor(actionName, key)` namespaced by action name only, with no
actor anywhere, so alice's stored `charge` response was returned to bob whenever
bob sent the same key; a differing payload gave bob X_IDEMPOTENCY_CONFLICT
instead, which is a cross-actor denial of service against any key. And
`Headers.get()` answers '' rather than null for `Idempotency-Key:`, so a blank
header was a live key every blank sender shared.

The status table is closed, so a code with no row falls to 500 — and stages.ts
reports every status >= 500 to the error monitor. Ten caller-caused codes were
in that state: a reused key, an expired cursor, a weak password, a duplicate
signup. The sharpest was the framework contradicting its own published
contract, action/http.ts:151 declaring '409' for X_IDEMPOTENCY_CONFLICT while
the runtime answered 500.

Nothing would have caught the eleventh, so the errors step gained a fourth host
rule. The specified predicate — every code owned by a tier <= 4 package needs a
row — flags 237 of 394, which is a step an agent disables in week one; and
whether a code can reach a request is not derivable, since
X_MIGRATION_DESTRUCTIVE and X_TENANCY_CROSS_DENIED are the same tier and the
same shape and blanket re-exports collapse import reachability to the whole
package. So it is a ratchet on the expectedRed idiom: 226 undecided codes
pinned with a reason, the list may only shrink, and a pin says "nobody has
decided yet" rather than "this can never reach a request". It caught two codes
on its first run, both added by its own teammates in this PR.

Also: a second server after a drain bound a port it could never serve from and
kept accepting connections after its own stop() returned; ?__proto__= replaced
the parsed query object's prototype, which also let a schema coerce an inherited
function through `key in record`; an empty issues array read as validation
success and reached a handler as an impossible undefined; the third copy of
FNV-1a/32 over client-chosen input, here backing both the idempotency
requestHash and the job dedupe key; and an unfenced settle overwriting a record
already replaced.

BREAKING CHANGE: `idempotencyKeyFor` takes a required third Actor argument; the
stored key's shape changed, so on the shared Postgres store a retry crossing the
deploy boundary re-runs the handler inside the 24h window (`truncate
x_idempotency` makes that state honest); `Idempotency-Key` is now enforced at the
255 characters OpenAPI already published; action's `fingerprint` is SHA-256/16,
changing job-handle.ts's dedupe key; `markReady()` throws X_LIFECYCLE_DRAINED on
a drained lifecycle instead of declining silently.

New codes: X_IDEMPOTENCY_KEY_INVALID, X_LIFECYCLE_DRAINED,
X_ERROR_STATUS_MISSING, X_ERROR_STATUS_BACKLOG_STALE, X_ERROR_STATUS_UNKNOWN_CODE.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RBwWKBJkiogA4mDaJiJf3D

* docs(plan): slices 02 and 06 are done across #107, #110 and #112

auth and entity needed nothing — every finding naming them was already closed by
#104 and #106, verified file-by-file before #112 was scoped. That is what took
the PR from an estimated ~70 files to 43.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RBwWKBJkiogA4mDaJiJf3D

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant