Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
197 changes: 197 additions & 0 deletions CHANGELOG.md

Large diffs are not rendered by default.

3 changes: 3 additions & 0 deletions biome.json
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,9 @@
"correctness": {
"noUnusedVariables": "error",
"noUnusedImports": "error"
},
"nursery": {
"noFloatingPromises": "error"
}
}
},
Expand Down
71 changes: 59 additions & 12 deletions docs/plans/2026/08/16/101-deep-dive-bug-audit/status.yml
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,12 @@ status: in_progress
created_by: sebi
worked_by: "claude (hive of 4 per PR, one checkout)"
owner: sebi
percent: 27
current_focus: "01+09+04 next (tiers 0-1, logic, projection). 05/10 = #101; 07 = #102/#103/#104, all merged"
percent: 60
current_focus: >
02+06 land their cache/query/jobs/core/db half as #110 — the realtime half was #107, so both
slices sit at 70% with only http, action, auth and entity left. Merged so far: 05/10 = #101;
07 = #102/#103/#104; 01/09 = #105; 04 = #106; 02/06 realtime = #107.
Next after #110: 02+06's http/action/auth/entity remainder, then 03.
slices:
- file: 05-gate-and-scripts.md
tier: 0
Expand All @@ -24,24 +28,29 @@ slices:
prs: [102, 103, 104]
- file: 01-tier01-bugs.md
tier: 0
status: not_started
percent: 0
status: done
percent: 100
pr: 105
- file: 09-logic-edge-cases.md
tier: 0
status: not_started
percent: 0
status: done
percent: 100
pr: 105
- file: 04-projection-contract.md
tier: 0
status: not_started
percent: 0
status: done
percent: 100
pr: 106
- file: 02-tier23-bugs.md
tier: 2
status: not_started
percent: 0
status: in_progress
percent: 70
prs: [107, 110]
- file: 06-concurrency-lifecycle.md
tier: 2
status: not_started
percent: 0
status: in_progress
percent: 70
prs: [107, 110]
- file: 03-tier45-bugs.md
tier: 4
status: not_started
Expand Down Expand Up @@ -91,6 +100,44 @@ evidence:
- "X_PACKAGE_UNREFERENCED fail-opens on a JSONC root tsconfig (comment => parse
returns undefined => rule reads it as 'no project references')."
new_codes: [X_REFERENCE_APP_NO_FLOOR, X_SCAFFOLD_GATE_RED, X_PACKAGE_UNREFERENCED]
- pr: 110
slices: [02, 06]
gate: "14 of 17 passed, 3 skipped (drift, contract-diff, budgets); app gate green, every pin
holds — examples/dummy 10/17 (7 red, 7 pinned), social-media-clone 14/17 (3 red, 3 pinned)"
falsified:
- "02-3b: 'reverse the Redis set ordering' is HALF a fix and worse alone. SADD-then-SET with
no re-check lets the bust SREM the membership and the later SET publish a row unreachable
by ANY tag — permanently uninvalidatable. Ships with an SISMEMBER re-check."
- "02-cursor: fixing `queryHash` alone was too narrow. `fingerprint` also backs
`cacheKeyFor`, so the 32-bit hash over client-chosen input was ALSO the shared cache
entry key — strictly worse than the cursor case the sweep named."
- "06-attempt: repeated suspensions cannot drive `attempt` negative. `nack` is fenced on
state === 'running', which only `claim` sets, and `claim` increments."
- "06-deadline: `configureLifecycle({deadlineMs})` is NOT unused — http/src/server.ts:97
declares one on every createServer. And X_SHUTDOWN_TIMEOUT already existed."
- "02-dates: the bare-Date.now() list in the cache package is EMPTY — fixed in #105/#106/#107.
Only `remember` bypassing assertTtl was real."
decided:
- "The drain deadline flipped from opt-in to bounded-by-default (25s). Built opt-in as
briefed, but `jobs` and `realtime` declare no budget, so the proven symptom — a worker pod
SIGKILLed mid-job — stayed unfixed. A mechanism that does not fix the symptom while
claiming the deadline works is a claim wider than what is enforced."
- "Colocated test fakes ship in the tarball. Four now do; a one-off `files` negation in one
package invents a convention nothing else follows. If it matters it is one sweep over all
29 package.json, and it belongs to the dead-code slice."
deferred:
- "The cache fill fence is PER PROCESS — two pods still interleave a load on one with a write
and bust on the other. Cross-node needs a Redis-side epoch and a wire change. Known-Gaps
row landed."
- "`createCacheStack` has zero production callers, so its copy of the fence is dormant. The
live fenced path is runQuery -> readRows -> readThrough -> fill."
- "`recentTierFailures()` has no reader outside cache, and /_x's invalidations source is
`unwired` — routed to slice 03."
- "`invalidateWireTags` has zero callers but is the wire-form door `x cache bust` will call.
Not dead; recorded for slice 15 to weigh."
- "`cache.scope` is auditable by grep, not by `x queries describe` — publishing it on
QueryDescriptor ripples into both tracked apps' manifests."
new_codes: [X_JOB_SLOT_LOST, X_QUERY_CACHE_TTL_INVALID]
notes: >
Slice order in this file is execution order, not tier order — 05 and 10 land first because the gate
does not yet mean what it says, and 08 lands late because it is mostly deletions that would collide
Expand Down
5 changes: 3 additions & 2 deletions dummy/social-media-clone/biome.json
Original file line number Diff line number Diff line change
@@ -1,11 +1,12 @@
{
"$schema": "https://biomejs.dev/schemas/2.4.15/schema.json",
"$schema": "https://biomejs.dev/schemas/2.5.5/schema.json",
"root": false,
"extends": "//",
"files": { "includes": ["**", "!x.manifest.json", "!openapi.json"] },
"formatter": { "indentStyle": "space", "indentWidth": 2, "lineWidth": 100 },
"linter": {
"rules": {
"recommended": true,
"preset": "recommended",
"suspicious": { "noExplicitAny": "error" },
"correctness": { "noUnusedVariables": "error", "noUnusedImports": "error" }
}
Expand Down
12 changes: 11 additions & 1 deletion framework.manifest.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

66 changes: 59 additions & 7 deletions packages/cache/CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,9 +14,32 @@ Tier 1. Tagged caching + THE invalidation graph.
`tier.invalidateTags()` from outside it. It is also the only place the log is written:
`recentInvalidations()` is a read of what that one path already reported, never a second
recorder a caller has to remember to call.
- **The fan-out clears FARTHEST tier first, and reports in read order.** Near-to-far leaves the far
tier holding the old value after the near ones are clear, and a read racing the bust promotes it
straight back up — `report.errors` empty, LRU stale again before the call returns. `CacheStack.drop`
reverses for the same reason. The report is re-sorted into `TIER_ORDER` because it is what the
`/_x` panel renders. Pinned in `invalidation-race.test.ts`.
- **A fill is fenced: sample before `load()`, ask before the write** (`fence.ts`). A read-through
fill publishes rows `load()` read in the past, so a bust landing in between finds a key that is
not there yet, reports `errors: []`, and is overwritten milliseconds later — invisible for the
whole TTL. `sampleFence({ key, tags })` → `fence.isValid()` is the whole API, `markInvalidated`
is its write half (called by `fanOut`, `CacheStack.write` and `CacheStack.drop`; a caller only
needs it for a clearing path of its own). It is **exported** because `@ultimat3/query`'s read
cache has the same hole and must not grow a second mechanism. `cover()` widens a fence
RETROACTIVELY — needed only where joiners contribute tags the leader never sampled, which is
`createCacheStack.read` and nothing else. A fence never fails a read: it declines to publish.
This is the one process-global here with **no `isolate*()` seam and no reset**, and that is
structural: a fence samples the current generation, which is always at or above what the ring has
forgotten, so another file's marks cannot invalidate a fence sampled after them.
- One graph. `graph.ts` exports functions over module state and **no constructor** — do not
add one, do not add a second registry anywhere else.
- Tag order is `TIER_ORDER`, never registration order. `sortTiers()` enforces it.
- **`bestEffort()` is public, and it is the only sanctioned way to swallow a cache refusal.** A
store outside this package that wraps its own `try/catch` degrades invisibly, and a second
failure log nobody reads is what this bounded one exists to prevent. Its label is `TierLabel` —
`TierName` plus `'query-read'` — closed, and deliberately NOT a widening of `TierName`: a name
missing from `TIER_ORDER` sorts to `-1`, ahead of the request memo. A label is a log facet; a
`TierName` is a position on the ladder.
- Tier failures go into `report.errors`. A cache tier may never fail a business read or write.
`createCacheStack` routes every `get`/`set`/`del` through `bestEffort()` for that reason — a
refusal becomes "that tier did not answer" and lands in `recentTierFailures()`, the read side's
Expand All @@ -43,7 +66,11 @@ Tier 1. Tagged caching + THE invalidation graph.
`isolateTiers()` already covers.
- Clocks are injected (`LruOptions.clock`, `CacheStackOptions.clock`); read them through `nowMs()`.
- **`ttlMs` is positive and finite, and `assertTtl` (in `tiers.ts`) is the one place that says so.**
Every tier calls it before it writes. `0` used to be "never expires" here and `EX 1` in `redis.ts`,
Every tier calls it before it writes, and so does `createMemorySemanticCache.remember` — which was
the one writer skipping it, so `ttlMs: 0` stored an entry already past its expiry and every lookup
missed with a completion bill as the only evidence. Its scope is `'semantic'` (`TtlScope`), with
`jitterFraction: 0`: spreading a lease is a herd defence for a SHARED store, and that one is per
process. `0` used to be "never expires" here and `EX 1` in `redis.ts`,
so one stack answered two ways; the rule lives beside `CacheSetOptions` precisely so a new tier
cannot invent a third reading. `X_CACHE_TTL_INVALID`, never a resolution.
- **`assertTtl` also SPREADS the lease it validated** — validate, then jitter, one choke point. A
Expand All @@ -56,6 +83,12 @@ Tier 1. Tagged caching + THE invalidation graph.
`realtime`'s `entry.reading`). The share ends as the load settles — a REJECTED load must clear
its entry too, or one origin failure becomes a permanent cached rejection. One `SingleFlight` per
stack, never a module-level map: two stacks are two ladders.
- **A joiner shares the leader's WRITE, so it contributes to it** (`FlightJoin`, merged by
`mergeSetOptions` in `set-options.ts`). Keyed on `key` alone and read late, the entry used to land
carrying only the leader's tags: the joiner's tag reached nothing, so the invalidation it declared
never fired. Tags union, TTLs take the SHORTEST — an entry held longer than a caller asked for is
stale to that caller. `work` reads the merge through `shared()` **after** the load, or it sees
only what the leader brought.
- **`negativeTtlMs` is the stack's decision, not a tier's.** Only `createCacheStack` sees what
`load()` answered, so the `null`/`undefined` branch lives in `ttlOptionsFor` there and reaches a
tier as an ordinary `ttlMs`.
Expand All @@ -68,11 +101,28 @@ Tier 1. Tagged caching + THE invalidation graph.
owns the clock, so it survives skew between the node that wrote and the node that reads, and no
stored payload shape changes under a running deployment. `-1`/`-2` are sentinels, not durations:
they mean no expiry, never one millisecond ago.
- **`redis.ts`'s script deletes only keys it was handed in `KEYS`.** The members of a tag set are
value keys in slots this node may not own, so `DEL`ing them from Lua is a cross-slot access that
fails on Redis Cluster and Dragonfly strict mode — into `report.errors`, so the bust reads as
partial and stale rows serve until TTL. The script returns the members; the tier deletes them
client-side, one key per `DEL`, which is slot-local under every topology.
- **`redis.ts`'s script deletes NOTHING — it reads.** The members of a tag set are value keys in
slots this node may not own, so `DEL`ing them from Lua is a cross-slot access that fails on Redis
Cluster and Dragonfly strict mode — into `report.errors`, so the bust reads as partial and stale
rows serve until TTL. The script returns the members; the tier deletes them client-side, one key
per `DEL`, which is slot-local under every topology.
- **The bucket is not dropped in the script either, and the tier `SREM`s only what it deleted.**
Dropping it atomically with the `SMEMBERS` made one failure permanent: a refused `DEL` left its
member with no bucket to be found in, so the retry the error asks for answered `keys: []` and
those rows served until their own TTL. `Promise.allSettled` is what makes "what actually died"
knowable. A member a concurrent write added between the two halves keeps its membership instead
of being orphaned by a bust that never deleted it.
- **A `set` joins its buckets BEFORE it writes the value, and re-checks membership after.** Value
first left a window where a bust's `SMEMBERS` saw an empty bucket and the value survived its own
invalidation for the full TTL. Joining first moves the window somewhere observable: membership
gone by the time the `SET` lands means this write was busted in the air, and the value goes with
it — a row nothing can reach by tag is one no later bust can clear. Only a literal `0` from
`SISMEMBER` counts as gone (`saysAbsent`); a reply the tier cannot read is not evidence, and
deleting on one is a cache that never caches.
- **`GET` and `PTTL` are two commands and the key can die between them.** `PTTL: -2` for a value
the `GET` returned is a MISS, not an entry with no expiry — reported as a hit it is promoted into
the LRU on the CALLER's ttl, so a row one millisecond from death gets a fresh five minutes one
tier closer.
- **That fixed half of it; `KEYS` itself was the other half.** Tag keys carry a `{entity}` hash tag
(`<ns>:t:{post}`, `<ns>:t:{post}:7`) and `invalidateTags` issues **one script call per tag**, so
every key a call is handed hashes to one slot. A single `EVAL` carrying two tags' buckets is
Expand Down Expand Up @@ -133,7 +183,9 @@ Tier 1. Tagged caching + THE invalidation graph.
|---|---|
| `tags.ts` | `tag` factory, wire form, match semantics, declared-tag registry |
| `graph.ts` | tag → dependents (cache keys, ISR routes, CDN paths, live queries) |
| `tiers.ts` | `CacheTier`, `TIER_ORDER`, read-through stack |
| `tiers.ts` | `CacheTier`, `TIER_ORDER`, `TierLabel`, `assertTtl`, read-through stack |
| `fence.ts` | the invalidation fence a fill (here or in `query`) checks before it publishes |
| `set-options.ts` | how two callers' `CacheSetOptions` combine, and the `null`-load TTL |
| `tier-failures.ts` | `bestEffort()`, and the bounded log of refusals it absorbs |
| `memo.ts` | request memo over the ALS ctx (WeakMap, no lifecycle) |
| `lru.ts` | byte-budgeted LRU (linked list + map + tag index) |
Expand Down
39 changes: 37 additions & 2 deletions packages/cache/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,7 +67,23 @@ a reader arriving while another's load is running joins it instead of issuing it
ends as the load settles, rejection included, so one failure is never held as a permanent one. A
feed cached for 60s and read 8,000×/s otherwise sends ~1,600 identical queries to Postgres at every
TTL boundary, because the write only lands after `load()` resolves. The primitive is
`createSingleFlight()` if you need it elsewhere; the stack holds one per stack.
`createSingleFlight()` if you need it elsewhere; the stack holds one per stack. A joiner shares the
leader's **write** as well as its load, so it contributes to it: tags union, TTLs take the shortest.
Without that the entry landed carrying only the leader's tags and the joiner's invalidation never
fired.

**A fill obeys an invalidation that raced it.** `load()` answers with rows it read in the past, so a
bust landing in between finds a key that is not there yet — it reports `errors: []` and the fill
republishes the pre-write rows for the full TTL, invisibly. `stack.read` samples a fence before the
load and re-checks it before each tier write; a fill that lost the race is dropped, and anything it
already wrote is taken back. The caller still gets what the origin answered: a fence declines to
publish, it never fails a read. It is exported for any cache doing its own read-through:

```ts
const fence = sampleFence({ key, tags });
const value = await run();
if (fence.isValid()) await tier.set(key, value, { tags });
```

**A `null` can carry its own TTL.** `negativeTtlMs` is used when the loaded value is `null` or
`undefined`, so a lookup for a row that has not replicated yet is not held for the positive lease:
Expand All @@ -93,6 +109,10 @@ though it were the value.
operation, the key and the `X_*` code, and each one also logged as `cache.tier.failed`. Same
bargain as `report.errors` on the invalidation side — degraded is visible, not merely slow.

`bestEffort(label, op, key, run)` is that guard, exported: a cache that is not a rung of this ladder
(`@ultimat3/query`'s read cache, `label: 'query-read'`) degrades into the same log rather than a
private `try/catch` nobody can read. The label is closed (`TierLabel`) so the panel can group by it.

## Tags

```ts
Expand Down Expand Up @@ -134,6 +154,13 @@ its collection's bucket hash to one slot, so a script may take both in `KEYS`. I
partial bust while stale rows served until TTL. Value keys are still deleted client-side, one `DEL`
each, which is slot-local under every topology.

The script **deletes nothing at all** — not the value keys, and not the buckets either. The tier
`SREM`s exactly the members whose `DEL` succeeded, so a refused delete keeps its membership and the
retry the error asks for still finds it; dropping the bucket inside the script made that failure
permanent. A `set` mirrors it: buckets are joined **before** the value is written and membership is
re-checked after, because a bust that landed in between would otherwise leave a row nothing can
reach by tag, serving until its own lease ran out.

**Every tag set carries a lease**, renewed on each write to the member's own TTL plus 60s, raised
only when the new lease is longer — a 60s member must not shorten a bucket a 1h member is in.
Without it a tag set grew forever: value keys died after five minutes, their membership never did,
Expand All @@ -157,7 +184,9 @@ your own payloads.
const report = await invalidateTags([tag('post', postId)]);
```

One function. Returns the report the `/_x` cache panel and `x cache bust --json` render:
One function. It returns the report below, which is also what the `/_x` cache panel renders — and
what `x cache bust --json` will print once it ships; that command is planned and exits
`X_NOT_IMPLEMENTED` today.

```json
{
Expand All @@ -174,6 +203,12 @@ One function. Returns the report the `/_x` cache panel and `x cache bust --json`
A dead tier lands in `errors` and never throws — a Redis outage must not fail the write
that triggered the bust. Entries there expire by TTL instead.

The fan-out walks the ladder **farthest tier first** — Redis before the LRU before the request memo
— and reports in read order. Clearing near-to-far leaves the far tier holding the old value after
the near ones are clear, and a read racing the bust promotes it straight back up into them: every
tier reports cleared and the LRU is stale again before the call returns. `stack.drop(key)` reverses
for the same reason.

`cdn` is what the dependency graph hangs off these tags, not what cleared: the `cdn` tier purges
those paths (as surrogate keys, alongside the tags), so what actually cleared is that tier's row
in `tiers`. With no `cdn` tier registered the list purges nowhere, which is why
Expand Down
Loading