Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -369,6 +369,10 @@ when in doubt, the code wins.

### Architecture docs (`docs/`) — describe the current system

- [system-contract.md](system-contract.md) — **normative**: the invariants and
binding rules the system guarantees (identity, read/write/cache contracts).
Where the descriptive docs and this contract disagree, the difference is a
bug. The read path implements it as of the #163 fix (PR #167). **Current.**
- [expression-lifecycle.md](expression-lifecycle.md) — one expression from MCP
ingest to rendered rows, naming every artifact and cache write. **Current.**
- [reactive-recalc.md](reactive-recalc.md) — revise an alias, recompute its
Expand Down
110 changes: 59 additions & 51 deletions docs/caching.md
Original file line number Diff line number Diff line change
Expand Up @@ -103,10 +103,13 @@ path), so there is no separate `result_cache/` directory; both land under
exempt (`_EXEMPT_READS`): re-reading a columnar source is already a pushdown.
- **Baked result snapshot** — an expression is *worthy* when it contains an
Aggregate / Join / Sort / window / UDF (`_EXPENSIVE_OPS`,
`result_cache.py:49`). A worthy expression is wrapped in a result cache with
`relative_path="result_cache"` (`source_cache.py:160`), so its snapshot lands
under `compute_cache/result_cache/`; executing the entry at build time
materializes it once (`build.py:491`). A non-parquet read is *not* by itself
`result_cache.py:49`). A worthy expression is first put in a canonical total
order (`_canonical_sorted`: the author's own `order_by` keys, then
`original_row_order`, then the remaining sortable columns — ADR D5, amended)
so the baked bytes are deterministic across the build and every later heal,
then wrapped in a result cache with `relative_path="result_cache"`, so its
snapshot lands under `compute_cache/result_cache/`; executing the entry at
build time materializes it once. A non-parquet read is *not* by itself
worthy (`classify_build`, `result_cache.py:62`) — its parse is already
handled by the source-read cache. Worthiness is decided by two predicates that
must agree: `_is_worthy_expr` gates the bake (`source_cache.py:62`) and
Expand All @@ -125,24 +128,33 @@ written for any entry, cheap or expensive, at build time or on demand (#73 and
follow-up).

Every consumer reads an entry's result through one function,
`cached_result_expr` (`result_cache.py:403`):
`cached_result_expr`, whose internals are the canonical read of
`docs/system-contract.md` (#163): expand the entry's frozen `xorq_build/`,
`load_expr(expanded, cache_dir=compute_cache)`, then the deep cache-dir
rewrite (`portable.rewrite_cache_dirs`). A missing or unloadable build is a
hard error naming the entry — reads never re-import `expr.py` (ADR D6):

- Expensive entry → `deferred_read_parquet` of the baked snapshot, on the
default backend. The snapshot was written when the entry built, so reading it
skips re-running the graph. If it was evicted since, the cache node recomputes
it once and the read proceeds — a self-heal re-checked on every call,
single-flighted per `(project, content_hash)` so concurrent cold readers don't
both run the shared cache op; after a repopulate the snapshot's digest is
checked against the recorded `result_digest` and a mismatch is logged as
recompute drift (`result_cache.py:451-473`, `_heal_lock`,
`_warn_if_self_heal_unfaithful`; #79, #83).
- Cheap entry → the live expression, re-imported from the entry's `expr.py`
and recomputed on read.
skips re-running the graph; the read asserts its derived snapshot key matches
the one the manifest recorded (ADR D8). If the file was evicted since, the
*frozen build* is executed once and the read proceeds — a self-heal
re-checked on every call, single-flighted per `(project, content_hash)` so
concurrent cold readers don't both run the shared op; the healed bytes are
verified against the recorded `result_digest` before they are served, and a
mismatch wipes the entry's Buckaroo stat cache, records a durable
`unfaithful_heal` error, and evicts its Buckaroo session (`_verify_self_heal`;
ADR D7/D10/D12, #79, #83).
- Cheap entry → the loaded build's graph, rebound onto the default backend and
recomputed on read from its content-pinned sources.

Either shape is a single-backend expression, so two entries compose (`union`,
`join`, a diff) without tripping xorq's "multiple backends" guard (#75). The
viewer's paginated reads, diffs, and post-processing all go through it; nothing
reads a pre-existing `result.parquet`, because none is written.
`join`, a diff) without tripping xorq's "multiple backends" guard. Chaining
(`tracked_expr_from_alias`) uses the sibling `entry_graph_expr` instead: the
parent's cache node stays in the graph, so a child's build is self-contained
and self-healing (ADR D4). The viewer's paginated reads, diffs, and
post-processing all go through `cached_result_expr`; nothing reads a
pre-existing `result.parquet`, because none is written.

Snapshot strategy is the right one here because an entry is an immutable,
content-addressed artifact — its result must not invalidate just because an
Expand Down Expand Up @@ -188,17 +200,17 @@ re-added entry still computes honestly cold. `reset_to` also reclaims orphaned
All bounded LRUs over immutable keys, so eviction means a cheap rebuild
and staleness is impossible:

- `cached_result_expr` — its reconstruction work is memoized on
`_resolve_result_plan`, `lru_cache(256)` keyed `(project, content_hash)`
(`result_cache.py:345`); `cached_result_expr` re-exposes that memo's
`cache_clear` / `cache_info`. The memo saves re-importing the entry's
`expr.py` and, for an expensive entry, re-deriving its baked-snapshot path.
`cached_result_expr` itself is a thin wrapper that re-checks the snapshot's
on-disk presence on every call, so an evicted snapshot self-heals; the heal
runs under a per-entry lock (`_heal_lock`, `result_cache.py:393`) so only the
first of several concurrent cold readers executes it, and a peer process that
wins xorq's fixed-`.tmp` rename is caught and retried once if the snapshot
landed, else re-raised (#79).
- `cached_result_expr` — its build-loading work is memoized on
`_resolve_result_plan`, `lru_cache(256)` keyed `(project, content_hash)`;
`cached_result_expr` re-exposes that memo's `cache_clear` / `cache_info`.
The key finally determines the value — the plan is a pure function of the
immutable build — and the memo saves expanding + loading the build and
re-deriving its baked-snapshot path. `cached_result_expr` itself is a thin
wrapper that re-checks the snapshot's on-disk presence on every call, so an
evicted snapshot self-heals; the heal runs under a per-entry lock
(`_heal_lock`) so only the first of several concurrent cold readers executes
it, and a peer process that wins xorq's fixed-`.tmp` rename is caught and
retried once if the snapshot landed, else re-raised (#79).
- `_build_compare_expr` — `lru_cache(128)` keyed
`(project, a_hash, b_hash, keys)`; saves rebuilding diff outer-join
expressions (`src/tallyman_companion/app.py:189`). Build dirs land under
Expand Down Expand Up @@ -336,23 +348,19 @@ all rely on "same key, same bytes, forever". Cleanup for those is a
space concern (manual delete, bullpen moves on reset), and a deleted file
self-heals by recomputing.

Two documented holes break "never stale". The first is execution
nondeterminism: an entry whose recipe calls `now()` / `random()` / an unseeded
`sample()` produces different bytes each run under one content hash, so a cold
One documented hole breaks "never stale": execution nondeterminism. An entry
whose recipe calls `now()` / `random()` / an unseeded `sample()` or an impure
UDF produces different bytes each run under one content hash, so a cold
recompute can disagree with what was built (#88). The build flags these as
advisory lint warnings (`_nondeterminism_warnings`, `build.py`); it does not
block them. As of #83 the build also records a `result_digest` of the executed
bytes in the manifest, and the self-heal compares the repopulated snapshot
against it (`verify_result_faithful`, `result_cache.py:299`;
`_warn_if_self_heal_unfaithful`, `result_cache.py:315`), so this drift is now
detected on recompute even though it is still not prevented. The second is
cold-reconstruction faithfulness under the `cas`
default: `cas` makes *build* identity content-aware (a rebuild over an edited
source forks the hash) but not *reconstruction*. `cached_result_expr` re-runs a
cheap entry's `expr.py` on every read, and `read_project_file` re-digests the live
source, so a cheap entry — or an expensive entry self-healing an evicted
snapshot — serves the edited bytes under its original `content_hash`. Closing
this needs digest-pinned reconstruction, tracked in #115.
block them. The runtime backstop is `result_digest`: every self-heal verifies
the repopulated snapshot against it before serving (`_verify_self_heal`), and
an unfaithful heal wipes the entry's stat cache, records a durable error, and
evicts its Buckaroo session (ADR D7/D10/D12). The second hole this section
used to document — cold reads re-running `expr.py` and re-digesting live
sources, serving edited bytes under the original hash (#115/#163) — is closed:
reads load the frozen build, whose leaves are content-pinned `.cas` clones and
inlined parent graphs, so a cold read cannot see a post-build edit at all.

Where staleness is actually possible, it is handled explicitly:

Expand All @@ -366,11 +374,11 @@ Where staleness is actually possible, it is handled explicitly:
Buckaroo restart, and the stat-cache deletion on reload.

The one rule to remember: snapshot-strategy caches do not notice upstream data
changes. The `cas` default papers over this at *build* time — an edited source
forks the entry hash rather than deduping to the stale one — but a cold read of
a cheap or evicted entry still re-digests the live source and can serve edited
bytes under the old hash (#74, above). Anyone who needs reconstruction to be
content-faithful today should treat a source as immutable once an entry is built.
`off` reverts to xorq's path-only identity (an edit collides silently); `salt`
mixes digests into the entry hash but bakes no snapshot cache at all (path-only
keys would collide across salted entries) and recomputes on every read.
changes. The `cas` default makes that safe end-to-end: an edited source forks
the entry hash at build rather than deduping to the stale one, and reads —
warm, cold, or self-healing — resolve through the frozen build to the
content-addressed clone the entry was built from, never the live file. `off`
reverts to xorq's path-only identity (an edit collides silently, and a loaded
build reads whatever bytes sit at the recorded path); `salt` mixes digests
into the entry hash but bakes no snapshot cache at all (path-only keys would
collide across salted entries) and recomputes on every read.
66 changes: 35 additions & 31 deletions docs/expression-lifecycle.md
Original file line number Diff line number Diff line change
Expand Up @@ -194,17 +194,13 @@ The SPA issues `GET /{project}/api/entry/{hash-or-alias}`
1. `_maybe_restart()` (revives a crashed subprocess, throttled); bail to
`None` if Buckaroo isn't running — the SPA then falls back to a
`pandas.to_html` preview.
2. **Snapshot self-heal** (`buckaroo_lifecycle.py:518-533`): before anything
Buckaroo-facing, `ensure_session` calls `cached_result_expr(project, hash)`.
For an expensive entry this guarantees the baked snapshot exists on disk — on
a cold compute cache (a fresh clone, or an entry the startup warmup didn't
reach) it recomputes it once, single-flighted per `(project, content_hash)`
under `_heal_lock` so two tabs hitting the same cold entry don't both run the
shared cache op (#79). This matters because a `tracked_expr_from_alias` chain off an
expensive parent embeds a *bare* read of the parent's snapshot (#75 strips
the parent's cache node to keep one backend), so without the snapshot the
replay below would read zero rows. Done outside the session lock so a cold
recompute doesn't block other sessions.
2. **No pre-heal** (ADR D4): entry builds are self-contained — a
`tracked_expr_from_alias` chain keeps the parent's cache node in the child's
build — so Buckaroo's replay of the posted build regenerates any evicted
snapshot itself through ordinary cache mechanics on first query. (Before
#163's fix, chaining stripped the parent's cache node to a bare snapshot
read, and `ensure_session` had to pre-execute `cached_result_expr` here so a
cold grid wasn't empty; both the stripping and the pre-heal are retired.)
3. **Session-map check** (`buckaroo_lifecycle.py:535`+): for a new entry this
misses, so a session is created.
4. **Stable build expansion**: `ensure_expanded_build` materializes
Expand Down Expand Up @@ -239,23 +235,27 @@ by `GET /{project}/api/session/{hash}`, `app.py:617`, which is just
row window and the summary-stats payload. The JS `SmartRowCache` holds row
segments client-side from here on (see `caching.md`).

The **result cache** is read through one function, `cached_result_expr`
(`result_cache.py:403`), and — unlike the stat cache — it was already populated
at build (step 5), so a read never writes it except to self-heal an eviction.
For an expensive entry (`classify_build` worthy — Aggregate / Join / Sort /
window / UDF, `result_cache.py:49`) it returns a `deferred_read_parquet` of the
baked snapshot under `compute_cache/result_cache/`, on the default backend; if
the snapshot was evicted it runs the cache node once to repopulate, then reads
(`result_cache.py:451-473`). That heal is single-flighted per `(project,
content_hash)` under `_heal_lock`, so concurrent cold readers don't both execute
the shared cache op and trip DataFusion's "Already borrowed", and a peer process
losing xorq's fixed-`.tmp` rename retries once if the snapshot landed (#79). The
repopulated snapshot is then checked against the `result_digest` recorded at
build, logging recompute drift on a mismatch (`_warn_if_self_heal_unfaithful`;
#83). A cheap entry returns the live expression re-imported from `expr.py`,
recomputed on read. Every reader takes this path: the paginated viewer
(`api_data` calls `cached_result_expr(...).limit(...).execute()`, `app.py:517`),
both sides of a diff, and post-processing.
The **result cache** is read through one function, `cached_result_expr`, whose
internals load the entry's frozen build (the canonical read of
`docs/system-contract.md`) — and, unlike the stat cache, it was already
populated at build (step 5), so a read never writes it except to self-heal an
eviction. For an expensive entry (`classify_build` worthy — Aggregate / Join /
Sort / window / UDF) it returns a `deferred_read_parquet` of the baked snapshot
under `compute_cache/result_cache/`, on the default backend, after asserting
the derived snapshot key matches the one the manifest recorded (ADR D8); if
the snapshot was evicted it executes the frozen build once to repopulate, then
reads. That heal is single-flighted per `(project, content_hash)` under
`_heal_lock`, so concurrent cold readers don't both execute the shared op and
trip DataFusion's "Already borrowed", and a peer process losing xorq's
fixed-`.tmp` rename retries once if the snapshot landed (#79). The healed
bytes are verified against the `result_digest` recorded at build before they
are served; a mismatch records a durable `unfaithful_heal` error, wipes the
entry's stat cache, and evicts its Buckaroo session (`_verify_self_heal`; ADR
D7/D10/D12). A cheap entry returns the loaded build's graph rebound onto the
default backend, recomputed on read from content-pinned sources. Every reader
takes this path: the paginated viewer (`api_data` calls
`cached_result_expr(...).limit(...).execute()`), both sides of a diff, and
post-processing.

---

Expand Down Expand Up @@ -332,6 +332,10 @@ Everything keyed on the content hash is immutable and self-heals on deletion;
the only invalidations are Buckaroo restart (RAM session map) and a
klass-changing edit/reset (stat cache). A `reset_to` additionally
garbage-collects orphaned `data/.cas` source clones, which live outside the
catalog git repo (`source_identity.gc_cas`). See `caching.md` for the full
invalidation table, including the cold-reconstruction staleness hole the `cas`
default does not close (#74/#115).
catalog git repo (`source_identity.gc_cas`) — an entry's own
`manifest.sources` records every clone its frozen build reads, ancestors'
included, so gc keeps a child's leaves alive even if the parent entry is
evicted. See `caching.md` for the full invalidation table. (The
cold-reconstruction staleness hole this section used to reference — #74/#115 —
is closed: reads load the frozen build, so a cold read cannot see a post-build
source edit.)
Loading
Loading