Skip to content

ADR-011 stage 1: a file enters the catalog only by an explicit import - #217

Merged
paddymul merged 31 commits into
feat/adr-007-009-cache-redesignfrom
feat/adr-011-sources-are-aliases
Sep 24, 2026
Merged

paddymul merged 31 commits into
feat/adr-007-009-cache-redesignfrom
feat/adr-011-sources-are-aliases

Conversation

@paddymul

@paddymul paddymul commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Stage 1 of plans/ADR-011-sources-are-aliases.md: the import path and the refusals (D1, D2, D3, D5, D9, D10, D12). D6 (deleting the staleness source axis), D8 (deleting the source-identity modes) and the rewrite of the 309 read_project_file call sites are stage 2.

A raw input stops being something a build reads. It becomes an alias whose versions are ordinary catalog entries.

What lands

  • update_and_depend(outside_path, alias, pinned_version=None, **reader_options) (src/tallyman_xorq/source_import.py) copies the bytes into the arena, writes ONE parquet snapshot of them in file order with __row_order at compute_cache/result_cache/<hash>.parquet, writes the entry, and appends a version to the source alias. Every row of D3's case table, plus D11: history is append-only, so no two version numbers ever denote the same bytes.
  • A source entry's content hash is md5 over its bytes and its reader options. build_expr cannot supply it — the generated recipe reads the snapshot and the snapshot is named by the hash. Two imports of one file under two aliases mint one entry and share one file.
  • The snapshot is data, not cache (ADR-007 D13): ensure_materialized refuses to re-create it and names the import that would; the Cache page's delete refuses it with the same reason.
  • D2 — read_project_file and tallyman_read_csv are build errors in an authored recipe, naming catalog_import_source. They survive inside the importer's generated recipe, where a contextvar resolves them to that entry's own snapshot.
  • D5 — pinned_expr_from_alias takes "<alias>-v<N>" only; a bare content hash gets an error naming the version reference for that hash.
  • D9 — ensure_cas_path digests the clone it wrote and refuses to publish one whose name lies about its content; recon_cas_path raises instead of serving drifted live bytes.
  • D10 — catalog_import_source replaces catalog_load_parquet: it takes a checkpoint, emits the events a revise emits, and cascades on the same per-project auto-recalc switch.
  • D12 — reader options are recorded on the entry at import and never re-derived.
  • A source alias is a distinct kind in the alias store, refused by catalog_create, catalog_revise, catalog_alias and by promoting a diff onto it; a name is one kind or the other, never both.

TDD

  • Red: c72d8bf, run 35779544514 — ruff green, fast suite 44 failed / 864 passed, all 44 new.
  • Green: ebfb971, run 35783080892 — ruff, fast suite and integration suite all green.

Deliberate debt

TALLYMAN_LEGACY_FILE_READS=1 disables D2's refusal and tests/conftest.py sets it for the whole suite, so stage 1 can land before the 309-call-site rewrite. Nothing in production sets it; the tests of D2 clear it per-test; it goes with the rewrite.

Judgement calls and what stage 2 now needs are recorded under Implementation notes in the ADR.

Based on feat/adr-007-009-cache-redesign, not main.

🤖 Generated with Claude Code

paddymul and others added 3 commits September 22, 2026 15:59
…ort (proposed)

A child that pins its parent by content hash is permanently stale on the
staleness scan's source axis and reports itself as an UNEXPLAINED orphan:
its recipe replays to the same hash, so the recorded source digest is never
refreshed and no user action clears the flag.

The cause is that manifest.sources is written as a retention record for the
.cas GC and read as a freshness record by the scan. Rather than teach the
scan the difference, this ADR removes the axis: a raw input becomes an alias
whose versions are ordinary entries, files enter only through an explicit
import, and a recipe names aliases and never a bare hash.

Twelve decisions, four of which delete mechanisms: the source axis and
manifest.sources (D6), the source-identity modes (D8), the ordered-copy key
and manifest.ordered_copies (D1), and the unfaithful live-bytes fallback in
recon_cas_path (D9).

Proposed. Nothing implemented.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tage 1)

The failing tests for the import path and its refusals. Nothing is
implemented yet; all 44 are red.

tests/test_source_import.py covers every row of the D3 case table for
update_and_depend, including the three errors and D11's "these bytes are
an older version" refusal; D1 (a source entry is worthy and its one
snapshot, named by the entry hash, IS the ordered copy, so nothing writes
compute_cache/ordered_sources/); the snapshot as data rather than cache
(ensure_materialized does not re-create it, the Cache page does not offer
to delete it); D12 (reader options are recorded at import and the same CSV
under two readers is two entries); D9 (a clone is digested after it is
written, and a lost version raises instead of serving drifted live bytes);
and D10 (catalog_import_source replaces catalog_load_parquet, takes a
checkpoint, records alias_set and triggers auto-recalc).

tests/test_io.py covers D2 (read_project_file and tallyman_read_csv are
build errors in an authored recipe, and the error names the import call,
while the importer's generated recipe still reads the file) and D5 (a bare
content hash in pinned_expr_from_alias is refused and the error names
"<alias>-v<N>").

tests/test_aliases.py covers the source alias as a distinct kind: recorded
in the store, a name is one kind or the other and never both, and a rename
keeps the kind.

The one other red test on this run, test_fouc.py::test_unknown_api_path_404s,
fails at bb28f80 too and is unrelated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tage 1)

Makes the 44 tests of c72d8bf pass. A raw input stops being something a
build reads: it becomes an alias whose versions are ordinary entries.

update_and_depend(outside_path, alias, pinned_version=None, **reader_options)
in the new tallyman_xorq/source_import.py copies the bytes into the arena,
writes ONE parquet snapshot of them in file order with __row_order at
compute_cache/result_cache/<hash>.parquet, writes the entry (a generated
recipe, a frozen build, a schema, a manifest) and appends a version to the
source alias. It implements every row of D3's case table plus D11's rule
that history is append-only, so no two version numbers ever denote the same
bytes.

A source entry's content hash is md5 over its bytes and its reader options,
because the recipe reads the snapshot and the snapshot is named by the hash
— build_expr cannot supply it. Two imports of one file under two aliases
therefore mint one entry and share one file. The snapshot is data and not
cache (ADR-007 D13): ensure_materialized refuses to re-create it, naming the
import that would, and the Cache page's delete refuses it with the same
reason.

D2: read_project_file and tallyman_read_csv are build errors in an authored
recipe, naming catalog_import_source. They survive inside the recipe the
importer generates, where a contextvar resolves them to that entry's own
snapshot. D5: pinned_expr_from_alias takes "<alias>-v<N>" only, and a bare
content hash gets an error naming the version reference for that hash.
D9: ensure_cas_path digests the clone it wrote and refuses to publish one
whose name lies about its content; recon_cas_path raises instead of serving
drifted live bytes. D10: catalog_import_source replaces catalog_load_parquet
— it takes a checkpoint, emits the events a revise emits, and cascades on the
same per-project auto-recalc switch. D12: reader options are recorded on the
entry at import and never re-derived.

A source alias is a distinct kind in the alias store, refused by
catalog_create, catalog_revise, catalog_alias and by promoting a diff onto
it, and a name is one kind or the other but never both.

Tests that changed with the behaviour rather than with a new assertion: the
pinned-parent tests now pin "<alias>-v1" instead of a hash, the MCP and
replay tests call catalog_import_source, and conftest sets
TALLYMAN_LEGACY_FILE_READS for the suite — scaffolding so stage 1 can land
before the 309-call-site rewrite of stage 2, and documented to be deleted
with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@paddymul
paddymul marked this pull request as ready for review September 22, 2026 20:58
paddymul and others added 7 commits September 22, 2026 17:00
The status line said both "Stage 1 implemented" and "Nothing here is
implemented". Say it once: accepted, stage 1 in #217, stage 2 is D6, D8 and
the call-site rewrite.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…(ADR-011 stage 1b)

The failing tests for the two decisions of 2026-09-22. Nothing is
implemented yet; all 14 are red.

Decision 1, in tests/test_source_import.py: a source snapshot is written
by pyarrow in the layout every computed snapshot has — parquet format
version 2.6, a page index, 122,880-row groups — instead of by polars,
which has no option for either (pola-rs/polars#12752 is open against
1.44.2, and 1.40.1 is installed, so no upgrade reaches it). polars keeps
the CSV, where it is the only reader that holds file row order and
applies the schema DSL of ADR-005, and its batches go to the same
writer: the import never collects the whole frame, and a parquet import
does not touch polars at all, so a date32, a time32 and a map come
through as the file has them (#197).

Decision 2, in tests/test_source_import.py and tests/test_cache_files.py:
a source snapshot is CACHE. The clone in data/.cas holds the imported
bytes and the entry records the reader options, so ensure_materialized
re-creates a deleted one and verifies it against the recorded
result_digest like any other, a child reads after result_cache/ is
emptied, and the Cache page lists the row unpinned and takes the delete.
Only a missing clone pins it, and then the error names data/.cas and the
re-import that repairs it.

Replaces the two tests of the old reading, that a source snapshot is
data nothing re-creates.

The one other red test on this run, test_fouc.py::test_unknown_api_path_404s,
fails at c492b20 too and is unrelated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…11 stage 1b)

Both patch polars out to prove the import does not need it, and then read
the file back through polars, so they failed on their own helper rather
than on the thing under test. They read it with pyarrow now. The CSV one
also asserts the file's writer, so that it discriminates the streaming
route into pyarrow from polars' sink (which collects internally) rather
than only the collect.

Still red, with the rest of 1621c95.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…-011 stage 1b)

The re-creation of a CSV source's snapshot has only the manifest's record
of the reader options, which went through JSON — the round trip that
#198 found a hole in — so it is worth its own case beside the parquet
one. Red: nothing re-creates a source snapshot yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… D1)

compute_cache/result_cache/ held two shapes of parquet file: the
computed snapshots that materialize writes, and the source snapshots
that an import wrote with polars. polars cannot write the first shape —
1.40.1 (installed) and 1.44.2 (latest) expose neither the parquet format
version nor a page-index option, and pola-rs/polars#12752 is still open
— so the import hands its Arrow data to the writer materialize already
uses instead.

materialize grows the two pieces that writer needs: row_groups, the
regrouping loop lifted unchanged out of _stream_to_parquet, and
write_pinned_parquet, which puts a stream of batches on disk under
_PARQUET_OPTIONS with row groups of a given size. A parquet source needs
no parser at all now (pq.ParquetFile.iter_batches), so it keeps the
types its file has, where polars rewrote a date32 as a timestamp, a
time32[ms] as a time64[ns] and a map as a list of structs (#197). A CSV
still goes through polars, the only reader that holds the file's row
order and applies the schema DSL and inference ladder of ADR-005, but
polars no longer writes it: _materialize_ordered takes the write step as
an argument, and the import passes one that streams collect_batches into
the same writer, so a source larger than memory imports the way a big
one should.

Measured on the 1.5M-row tests/big_parquet.py fixture: format version
1.0 -> 2.6, created_by Polars -> parquet-cpp-arrow 21.0.0, the same 13
row groups of 122,880 rows, and 16,188,593 -> 25,495,419 bytes. The
growth is pyarrow's default dictionary encoding on a 122,880-row chunk
of int64 (~983 KB, just under its 1 MB dictionary_pagesize_limit, so it
never falls back to PLAIN); the same data is 16,205,749 bytes with
use_dictionary=False and 17,827,724 with 1,048,576-row groups. Whether
to change either is a question about _PARQUET_OPTIONS for every snapshot
at once (ADR-009), not a reason to write source snapshots differently
from computed ones. The page index turns out to be present either way:
polars writes one unasked.

Makes 6 of the 15 tests of 1621c95, d65222b and 1d40db0 pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Stage 1 kept the imported bytes in data/.cas/<digest><suffix> and then
treated the source snapshot as data that ensure_materialized must never
rebuild. Those two do not sit together: ADR-007 D13 says a file is cache
exactly when ensure_materialized can re-create it, and the clone plus
the reader options on the entry (D12) are everything needed to write the
snapshot again. What stage 1 actually shipped was a deleted
result_cache/ leaving every source entry unreadable with its data on
disk a directory away.

ensure_materialized now writes it again. It cannot go through
materialize — a source entry's build reads the very snapshot that is
missing — so _heal_a_source parses the clone through the import's own
writer and hands the digest to _verify_self_heal, the same check and the
same unfaithful_heal record every other re-created snapshot gets.
pinned_reason follows: a source row is ordinary deletable cache, and it
is pinned only when the clone is gone as well, which is the one case
where the snapshot is the last copy of those rows. A read of an entry
whose snapshot AND clone are both gone still raises, and the error now
names the clone it looked for as well as the re-import.

Makes the other 9 tests of 1621c95 and 1d40db0 pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
D1 said the source snapshot was data that nothing could re-create; the
clone the same stage kept makes that false, so the paragraph is rewritten
around what ADR-007 D13 actually asks, and the one case that is still a
loss — the clone gone too — is written down beside it.

The implementation notes record the polars finding so nobody re-opens
it: 1.40.1 installed, 1.44.2 latest, neither can set the parquet format
version or ask for a page index (pola-rs/polars#12752), and use_pyarrow
is pyarrow. They also record the two measurements that came out of the
change — the page index was there all along, and the pinned layout costs
57% on the big fixture for a reason that belongs to ADR-009 —
and why SNAPSHOT_FORMAT_VERSION stays 1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@paddymul

Copy link
Copy Markdown
Contributor Author

Stage 1b: the snapshot writer, and the snapshot as cache

Two decisions Paddy made after reading the stage 1 report, both on this branch.

pyarrow writes a source snapshot, in the pinned layout. Stage 1 wrote it with polars, so result_cache/ held two parquet shapes: computed snapshots with version="2.6" and a page index, source snapshots without. I checked whether a newer polars could write them itself before adding anything — installed is 1.40.1, latest on PyPI is 1.44.2, and the parameter list is identical: no page index, no format version. The open request is pola-rs/polars#12752. So polars stays only where it is needed (parsing a CSV, since it is the only reader that preserves file row order and applies ADR-005's schema DSL) and hands its Arrow output to the existing pyarrow writer. A parquet source skips polars entirely. One write, not a write plus a re-encode.

Verified by reading a source snapshot's own footer rather than trusting the options dict:

format_version  2.6
created_by      parquet-cpp-arrow version 21.0.0
has_column_index = True
has_offset_index = True
columns         [order_id, region, category, price, qty, __row_order]

A source snapshot is cache, re-created from the clone. Stage 1 treated it as data that ensure_materialized must never rebuild, which contradicts ADR-007 D13 — a file is cache only if it can be re-created, and the .cas clone holds the imported bytes, so it can be. A deleted result_cache/ left source entries unreadable with the data sitting right there. Now it is re-created from the clone with the recorded reader options and verified against the recorded digest. Only a missing clone and a missing snapshot is unrecoverable, and that is a loud error naming re-import.

This matches what #215 does for ordered copies (_row_groups / _write_parquet_copy). The two will need reconciling when #215 lands on the #189 branch; the approach is deliberately the same so the rebase is mechanical.

CI: red on 1621c95, d65222b and 1d40db0, green on ebe0702 (run 35808338219). Locally 913 passed plus the pre-existing test_fouc.py::test_unknown_api_path_404s, which needs the built SPA, and 7 integration.

Stage 2 — the 309 call sites, D6 and D8 — is on feat/adr-011-stage-2, stacked on this branch.

The assertions stage 2 has to make true, ahead of the code that makes them
true. A scan of a project with three imported sources and two children takes
no digest at all and reports no axis it could not evaluate; a manifest records
neither sources nor ordered copies; source_identity has no mode, no salt and
no stat-memoized digest, and no module names TALLYMAN_SOURCE_IDENTITY; the
ordered-copy store and its copy key are gone, so compute_cache holds one kind
of file; io folds no parent source records; the rebuild script replays a
source entry through an import rather than through build_and_persist; and
TALLYMAN_LEGACY_FILE_READS buys an authored recipe nothing.

The digest counter is its own positive control: it counts the two digests each
import takes (the file as given, then the clone as written) and zero across
the scan, so the claim is measured rather than read off the code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
paddymul and others added 7 commits September 22, 2026 22:37
update_and_depend(path, alias, project=X) fails whenever X is not the active
project. The recipe the importer generates calls read_project_file with no
project argument (source_import.py:348), so the read resolves the ambient
active project, the _SOURCE_ENTRY contextvar it is checked against names X,
and the two disagree — the importer's own recipe then hits the D2 refusal
written for authored recipes and the import dies with an error about a file
tallyman does not own.

Every call site in the repo passes today because the project fixture also
makes its project active. Rewriting the suite onto imports found it: the
cache lab warms xorq in a project of its own, without activating it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…11 D6, D8)

The source axis, manifest.sources, manifest.ordered_copies, the ordered-copy
store and the source-identity modes are all deleted. read_project_file and
tallyman_read_csv now always refuse an authored recipe, and the
TALLYMAN_LEGACY_FILE_READS escape hatch that let the suite keep authoring raw
reads is gone with them.

The test rewrite onto imported source aliases is PARTIAL: about half the suite
is moved over and the rest still calls the deleted helpers. Committed as-is
because four agents were editing this worktree at once and one of them lost
eight files to a git restore. Follow-up commits finish it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
update_and_depend(path, alias, project=X) failed whenever X was not the active
project. The recipe the importer generates calls read_project_file with no
project argument, so the read resolved the ambient active project while the
contextvar naming the entry being minted said X. The two disagreed, the
importer's own recipe hit the D2 refusal written for authored recipes, and the
import died complaining about a file tallyman does not own.

project= is an override, so the read now follows the entry being minted rather
than whichever project happens to be active. source_entry_context returns the
(project, hash) pair it already held instead of filtering on a project passed
in, and neither entry point resolves the ambient project any more.

Every call site in the repo passed before this because the project fixture also
activates its project. Rewriting the suite onto imports is what exposed it: the
cache lab warms xorq in a project of its own without activating it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The call-site rewrite left four files still opening a file in a recipe. Each
now imports it as a source alias and reads the alias.

Two of the four failures were the new behaviour, not breakage:

- the page-load profiler measures three entries where it measured two, because
  an imported source version IS an entry (D1). The corpus fixture says so
  rather than the assertion excusing it.
- the reset scenario has to import extra.parquet AFTER s1. Importing it before
  keeps its source entry alive across the reset back, so no clone is left
  referenced only by a retired entry and the case ADR-007 D13 is about stops
  being exercised. With the import after s1 the test proves what it claims: a
  backward reset parks the clone in the bullpen instead of unlinking it, and a
  forward reset restores it.

test_replay's storyboard imported the parquet and then authored a raw read of
it anyway; the stdio round trip did the same over the real MCP transport.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Each of the three ADRs that described how a raw input reaches a build now says
what ADR-011 changed, in its status block, so a reader who starts at any of
them is not misled:

- ADR-002 keeps its clone store and loses the rest: the mode switch and
  salted_hash, manifest.sources, digest_for and its stat memo, and the
  reconstruction caveat, which is now an error rather than a fallback to live
  bytes.
- ADR-005's reader runs at import, not in a recipe. The schema DSL, the
  inference ladder and the suggestion contract are unchanged; they just live
  inside catalog_import_source now, and the options are fixed per alias.
- ADR-008's refusal of a raw read extends to read_project_file itself, and the
  ordered copy stops being a separate kind of file: it is the source entry's
  snapshot, named by its content hash.

ADR-011 records stage 2 in its implementation notes, marks itself implemented,
and answers its first open question: keep the raw bytes, which is also what
makes the snapshot cache rather than data.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Same bytes mint the same entry hash in both projects, but each holds its
own entry, snapshot and clone: a heal, a clone sweep or a re-import in one
never reaches the other, and the one-alias-per-bytes rule is per project.

Also pins the way to give a source a second name: a catalog entry whose
recipe reads it. Its hash is xorq's hash of the expression, not the md5 a
source entry is named by, so the two never coincide and nothing has to be
added to the recipe to force them apart.

These pass today; they guard the duplicate-import fix that follows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A second import of bytes another alias already holds minted nothing new
but re-ran _mint over the shared entry: it rewrote the first alias's
manifest and recipe with the second alias's name, and a failure in the
build step rmtree'd the entry the first alias points at.

- the second import raises, naming the alias and version that hold the
  bytes and the catalog_create way to a second name, and leaves the first
  entry byte-for-byte as it was
- the same holds for bytes of an older version of another alias
- catalog_import_source returns it as an error and records no revision
- a re-import that repairs a missing snapshot keeps the existing entry
  when the build step fails

Replaces test_two_aliases_over_identical_bytes_share_one_entry, which
pinned the behaviour being removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
paddymul and others added 8 commits September 24, 2026 05:57
A source entry is named by its bytes and reader options, so a second
import of the same bytes under another alias landed on the first alias's
entry and re-ran _mint over it: the manifest's provenance and the recipe
were rewritten with the second alias's name, and a failure in the build
step rmtree'd the entry the first alias points at.

- update_and_depend refuses a mint whose entry is a version of another
  alias of the project, at any version. The error names that alias and
  version and the way to a second name: a catalog entry whose recipe is
  tracked_expr_from_alias(<that alias>). Keyed on the entry hash, so a CSV
  read two ways is still two imports (D12); per project.
- _mint removes the entry directory on failure only when it created it,
  so a re-import repairing a missing snapshot keeps the existing entry.
- ADR-011 D1 now says one set of bytes is one version under one alias,
  and records that it first said the opposite and why that changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The kind lives on the alias ("catalog" or "source" in aliases.jsonl), but
whether an entry is a source entry lives on the entry (its manifest carries
provenance), and nothing ties the two together. catalog_alias checks only the
kind of the name, so catalog_alias(<source entry hash>, "x") makes a catalog
alias whose head is an imported file. catalog_revise("x", ...) is then
allowed, which bypasses D1's "a source version has no recipe to revise".

Three tests pin the invariant:

- set_alias refuses a catalog alias onto a source entry;
- set_alias refuses a source alias onto a computed entry;
- MCP catalog_alias on a source entry returns an error naming the source
  alias-version (orders-v1) and steering to catalog_create over
  tracked_expr_from_alias, creates no alias, and the steer then works.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
manifest.provenance.alias and .version record the name a source version
was imported under, once. A rename carries the alias's history and kind to
the new name, and an unalias drops it, but three places still read the
import-time name as the current one:

- materialize.pinned_reason and the _heal_a_source error name the version
  a_src-v1 after a rename to renamed_src, and the heal error advises
  catalog_import_source(path, 'a_src', pinned_version=1). Following that
  advice mints a second source alias, a_src, over the same entry.
- The generated recipe's header reads "# a_src-v1: a source version", so the
  Code tab of renamed_src-v1 presents the old name as the entry's.
- scripts/rebuild_native_catalog.py replays a source entry under
  provenance["alias"] and orders its versions by that alias's history. After
  a rename the replay mints a_src, never makes renamed_src before the
  children build, and a child reading renamed_src fails; the versions lose
  their ordering edge and fall into hash order.

These tests pin the current name in both messages, advice that repairs the
renamed alias without minting a new one (through update_and_depend and
through the MCP tools), wording that does not send an unaliased version back
under its dead name, a recipe header that records the import name as
history, and a rebuild that replays a renamed source under its current name
and in version order.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
set_alias now reads the manifest of the entry it is pointing at. A manifest
that records provenance is a source entry, a version of an imported file; any
other manifest is a computed entry. A catalog alias onto a source entry, or a
source alias onto a computed entry, raises AliasKindMismatch. The check lives
in set_alias because every route that names an entry goes through it, so
catalog_alias, a revise, a promoted diff, a recalc and an import all keep an
alias's kind and its entries' kind the same without each tool repeating it.

A hash with no readable manifest is not checked: the alias bookkeeping is used
and tested with hashes that name no entry. A passing test pins that boundary.

For a source entry the message names the source alias-version that holds it
(via version_of_hash) and gives the ADR-011 way to a second name:
catalog_create('<name>', "...expr = tracked_expr_from_alias('<source>')"), a
catalog entry that reads the source and follows it when it is imported again.
MCP catalog_alias returns that message as {"error": ...}.

The other set_alias callers only alias a hash they just built (catalog_create,
catalog_revise, promote_diff in MCP and the companion, the companion's PUT
/api/code, recalc, the rebuild and perf scripts) or just imported
(update_and_depend), so the new check cannot fire for them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
provenance.alias and .version stay what they are: the record of the import.
Everything that names a source version as it is now asks the alias store.

- source_import.source_version_in finds the source alias whose history holds
  an entry, trying the name it was imported as first. It is
  aliases.version_of_hash restricted to source aliases, because an import
  advances nothing else. current_source_version applies it to a project.
- materialize.pinned_reason and the _heal_a_source error name the version by
  that alias and add "imported as <old>-vN" when a rename changed it. The
  advised catalog_import_source(...) names that alias and version, so
  running it repairs the snapshot and mints nothing. When no source alias
  holds the entry, both messages say so, and the error says an import of the
  same bytes writes them again as a new version of whichever alias it names.
- The generated recipe's first line reads "Generated by catalog_import_source
  when the file was imported as a_src-v1", which stays true after a rename.
- scripts/rebuild_native_catalog.py orders a source's versions and replays
  each one under the source alias that holds it in the old catalog, pinned to
  its version there, so an out-of-order replay is an error. The re-point loop
  now skips only the alias the import pointed and sets any other holder with
  the kind it had; the pin needs a second source alias over the same bytes to
  be on that version before its next one replays.
- SourceProvenance's docstring says alias and version are the import name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… routes refuse source aliases (ADR-011)

Found reviewing #219.

- A re-import of a version whose snapshot is gone went through _mint: it
  rebuilt the entry, rewrote the manifest (provenance path, imported_at,
  result_digest, row_count) and took whatever it wrote as the new truth.
  A probe with a reader that drops a row showed 10 -> 9 rows under the
  same hash and no unfaithful_heal record, where ensure_materialized
  records one. The repair should rewrite nothing but the snapshot, and be
  verified like any heal.
- An entry directory with no manifest (a crash part-way through an
  import) could not be repaired by re-importing: the no-op branch read
  the manifest and raised.
- The heal error's advised catalog_import_source(...) dropped a CSV's
  schema and reader options, so running it as written was refused as
  "not orders-v1". The pinned-version refusal also blamed the file when
  only the reader options differed.
- PUT /api/code/<source alias> and POST /api/promote_diff onto a target
  name that is a source alias built an entry, failed to point the source
  alias at it and answered 500. Both should refuse before building, 409.

Replaces test_a_failed_repair_leaves_the_existing_entry_in_place, which
pinned the repair going through _mint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
paddymul and others added 2 commits September 24, 2026 07:42
ADR-011 review fixes: one alias per bytes, alias kinds, current names, verified repairs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ADR-011 stage 2: one staleness axis, one identity mode, and no raw file reads

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@paddymul
paddymul merged commit 8530876 into feat/adr-007-009-cache-redesign Sep 24, 2026
paddymul added a commit that referenced this pull request Sep 24, 2026
ADR-011 held back docs/architecture.md, caching.md, expression-lifecycle.md and
system-contract.md until this PR landed, so they would be rewritten once. Now
that the PR is rebased onto the merged ADR-011 stack (#217, #218, #219), they
describe what exists: a file enters only by catalog_import_source, as a source
entry under a source alias; a source entry's hash is an md5 of its bytes and
reader options; its snapshot is cache, healed from the clone under data/.cas,
and pinned once the clone is gone; recipes name aliases, never files or bare
hashes; alias kinds and the rule that an alias's kind matches its entries;
one set of bytes is one version under one alias; and staleness has one axis.

The ordered copy, the identity modes, manifest.sources, the source-digest
memo and the source axis are gone from every doc except where one says what
was removed. Known defects #197, #198, #207, #211 and #191, and the two
staleness defects with no issue, move to a note that they no longer apply.
reactive-recalc.md's walk-through now imports orders and advances it by a
re-import. mcp-server.md, installing.md (TALLYMAN_SOURCE_IDENTITY is gone), the
README and the code-derived sections of tallyman_explanation.md follow suit.

The #193 to #196 passages are unchanged; those fixes are being ported onto
this branch separately.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
paddymul added a commit that referenced this pull request Sep 24, 2026
ADR-011 held back docs/architecture.md, caching.md, expression-lifecycle.md and
system-contract.md until this PR landed, so they would be rewritten once. Now
that the PR is rebased onto the merged ADR-011 stack (#217, #218, #219), they
describe what exists: a file enters only by catalog_import_source, as a source
entry under a source alias; a source entry's hash is an md5 of its bytes and
reader options; its snapshot is cache, healed from the clone under data/.cas,
and pinned once the clone is gone; recipes name aliases, never files or bare
hashes; alias kinds and the rule that an alias's kind matches its entries;
one set of bytes is one version under one alias; and staleness has one axis.

The ordered copy, the identity modes, manifest.sources, the source-digest
memo and the source axis are gone from every doc except where one says what
was removed. Known defects #197, #198, #207, #211 and #191, and the two
staleness defects with no issue, move to a note that they no longer apply.
reactive-recalc.md's walk-through now imports orders and advances it by a
re-import. mcp-server.md, installing.md (TALLYMAN_SOURCE_IDENTITY is gone), the
README and the code-derived sections of tallyman_explanation.md follow suit.

The #193 to #196 passages are unchanged; those fixes are being ported onto
this branch separately.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant