Skip to content

fix(xorq): an ordered copy keeps its source's types and reader options, and its digest file is written atomically (#197, #198, #211) - #215

Closed
paddymul wants to merge 4 commits into
feat/adr-007-009-cache-redesignfrom
fix/197-198-211-ordered-copies
Closed

paddymul wants to merge 4 commits into
feat/adr-007-009-cache-redesignfrom
fix/197-198-211-ordered-copies

Conversation

@paddymul

@paddymul paddymul commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #189: the base is feat/adr-007-009-cache-redesign, not main. Fixes #197, Fixes #198, Fixes #211.

An ordered copy is the parquet copy of a source that a recipe reads instead of the source: the source's rows in file order plus a last column, __row_order, 0..N-1 (ADR-008 D2, every file tallyman reads carries __row_order). It lives at compute_cache/ordered_sources/<key>.parquet, and a <key>.digest file beside it caches its content digest. The manifest records the reader options it was made with, so ensure_materialized can make a deleted copy again (ADR-007 D13, which files are cache).

What changed

Docs: ADR-008's and ADR-009's "Implementation notes", and the lines in docs/caching.md and docs/system-contract.md that said polars writes every copy.

The snapshot format version stays at 1

SNAPSHOT_FORMAT_VERSION (ADR-009 D3, the snapshot format) stands for the settings that decide the batch boundaries an entry sees, and the ordered copy's 122,880-row groups are one of them. The pyarrow copy's bytes differ from the polars copy's, but its row groups are where they were, and polars also wrote a page index. I measured the 1.5M-row test source (tests/big_parquet.py) on the single-partition materialization connection. The two copies have the same content digest. The SUM and AVG of the float column come out bit for bit the same ungrouped, filtered on the float, and over four ranges of the sorted id column, which the page index prunes. The two copies of orders.parquet also have matching digests. A version bump would make any later unfaithful heal of an existing entry blame a format change that did not change what the entry read.

What does change is the columns polars did not keep. An entry built before this over a date64, map, time32 or time64 column read a copy with other types. If its copy is made again, recreate_ordered_copy finds that the content digest differs from the recorded one and records unfaithful_ordered_copy. Rebuilding such an entry is the fix.

How it was built

  1. 2665f80: six failing tests. CI run: ruff passed and the fast suite failed on exactly those six (6 failed, 864 passed), each for the reason its issue gives.
  2. 1ac8fc7, 07eb3d0, c267f7d: one fix per issue, each making its own tests pass. CI run: ruff, the fast suite (870 passed) and the integration suite (7 passed) all pass.

Locally, on c267f7d: the fast suite gives 863 passed, 6 skipped and 1 failed. The failure is test_fouc.py::test_unknown_api_path_404s, which gets a 503 because this worktree has no packages/app/dist, and it fails the same way without these commits. I also checked the new writer on an empty source, a source whose only column is __row_order, a pandas categorical spanning row groups with different dictionaries, and list and struct columns.

Not in this PR

  • fixed_size_binary: xorq 0.3.26 cannot read it (KeyError: FixedSizeBinaryType while it builds the table's schema), from the source or from the copy. polars turned it into binary, so a recipe over one used to build. Now it fails the way a direct read of the source does.
  • The copy of a source with unique numeric columns is larger: 25.5 MB against 16.2 MB from polars for the 1.5M-row test source. The snapshot settings dictionary-encode every column, and an int64 or float64 dictionary never overflows in a 122,880-row group. Snapshots pay the same cost. Changing use_dictionary would change both files, so I left it.
  • The copy's key includes the JSON form of the reader options, and for a function that is its repr, which holds the function's memory address. So two builds of the same recipe with a function option share an entry only if the function lands at the same address. Four builds, two in each of two processes, gave four entries. Within one process the address is never reused, because each build's recipe module stays in sys.modules.
  • _row_groups repeats the regrouping in materialize._stream_to_parquet. I left materialize.py alone because fix(xorq): a failed build keeps the snapshot already on disk for its hash #213 edits it.

🤖 Generated with Claude Code

paddymul and others added 4 commits September 22, 2026 12:39
- #197: the ordered copy of a parquet source with date64, map, time32 and time64 columns has the source's schema plus
  __row_order, and a recipe sees the ibis types a direct read of the source gives. A decimal256 source builds. A
  polars panic while reading a CSV is a BuildError.
- #198: a function passed as with_column_names reaches polars as a function, and a deleted copy of that read cannot
  be made again from the manifest and says so.
- #211: an empty .digest sidecar is recomputed from the copy, not recorded as the copy's content digest.

The polars PanicException is a BaseException, so the tests that can hit it catch it and fail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… the source's types (#197)

polars wrote the copy and did not keep every Arrow type: a date64 came back as a timestamp, a map as a list of
structs, time32 and time64 as time64[ns], and a decimal256 made it panic with a PanicException, which no
`except Exception` catches. _write_parquet_copy now streams ParquetFile.iter_batches() in file order, regroups the
rows into 122,880-row groups, appends __row_order from np.arange and writes with the snapshot's parquet settings. The
copy's schema is the source's plus __row_order.

SNAPSHOT_FORMAT_VERSION stays at 1. The row groups are where they were, and on ordinary columns the new copy has the
same content digest and gives bit-identical float SUM and AVG on the single-partition connection; ADR-009's
implementation notes have the measurements.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…lars panic is a BuildError (#198, #197)

tallyman_read_csv built the copy from the JSON form of its schema and scan_csv options on the first ingest, the form
the manifest records. JSON turns a function into its repr, so with_column_names=lambda ... reached polars as a string
and polars panicked calling it. ensure_ordered_copy now takes the caller's (schema, scan_kwargs) as csv_args and the
first write uses them; the JSON form still names the copy and is what recreate_ordered_copy replays, which already
refuses a record marked lossless: False.

polars still parses a CSV, and a function it calls can still raise, which makes it panic. _write_copy turns that
PanicException, a BaseException that the build's `except Exception` and the MCP tool's `except BuildError` both
miss, into a BuildError that names the source.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… empty one is read back from the copy (#211)

The sidecar that caches a copy's content digest was written with Path.write_text, which truncates before it writes,
and was read back as whatever it held. A build that read it empty recorded content_digest "" in its manifest, and a
later recreate_ordered_copy then reported the copy as unfaithful.

_write_atomically now computes the digest from the temp file, writes the sidecar to a unique temp name and moves it
into place with os.replace, then moves the copy into place, so a reader that finds the copy finds its whole digest
beside it and never one left by an earlier copy. _content_digest_of treats an empty sidecar as missing: it reads the
digest back from the copy and writes the sidecar the same atomic way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@paddymul

Copy link
Copy Markdown
Contributor Author

Superseded by the ADR-011 stack (#217, #218, #219). ADR-011 deleted the ordered copy; the #197 fix (pyarrow writes a parquet source's snapshot) and the #198 fix (reader options fixed at import) are in the stack, and the .digest sidecar #211 was about no longer exists.

@paddymul paddymul closed this Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant