Repository navigation
fix(import): a failed import is recorded, names the user's file, and a column xorq cannot read is left out and named (#224, #225, #227, #234, #239) - #247
Conversation
Red for #224, #225, #227, #234 and #239, all in the ADR-011 import path. - #224: a parquet file with a fixed_size_binary column, a struct-nested one and a UUID column raises SourceImportError naming all three, before the clone or the snapshot is written. - #225: an import that fails in its generated recipe, and a CSV that fails to parse, leave result_cache/ and data/.cas/ as they were. - #227: through fastmcp's client, each CSV reader failure the issue lists, and a clone that fails its digest check, returns {error, error_id} with one errors.jsonl record, a build_error event and a build_failed notification. The message names the user's file and catalog_import_source, not the clone or tallyman_read_csv, and a parse failure ends with the retry in the tool's argument shape, which runs as written. - #234: the append-only refusal names tallyman revisions and tallyman reset-to <step>, and every catalog_* name in a message under src/ is a registered tool or a Python name. - #239: a re-import with the snapshot present restores a lost clone, at the path the entry names, and lifts the pin. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… before writing anything #224. The generated recipe of a source entry reads its snapshot with deferred_read_parquet, which asks xorq's DataFusion backend for the file's arrow schema and converts each field with PyArrowType.to_ibis. That map has no entry for fixed_size_binary, so a fixed_size_binary or UUID column failed there with a KeyError naming no column, after the clone and the snapshot were written. update_and_depend now registers a parquet file with the same backend, takes the arrow schema it reports, and runs the same conversion field by field before it digests or writes anything. The SourceImportError names every column and nested field that fails (meta.key, tags[]), with a cast: binary, string for a UUID, the storage type for an extension type. The check asks DataFusion, not pq.read_schema as the issue suggested, because the two schemas differ where it matters: DataFusion gives a top-level extension column its storage type, so a JSON, bool8 or fixed_shape_tensor column imports, and keeps a nested one, which the read then refuses. test_the_type_check_passes_every_type_the_read_takes pins the first half; it passes before and after. Nothing is executed, and the check costs about 6 ms per import. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eaves nothing it wrote #225: _mint notes which of the clone and the snapshot were absent when it started, and one failure handler removes those and the entry directory it made. A clone already there, which a CSV read another way shares, and a snapshot a reset left stay. test_a_failed_import_keeps_a_clone_another_entry_uses pins the first; it passes before and after. #227: a reader failure while the snapshot is written becomes a SourceImportError that names the outside file and catalog_import_source, and CloneDigestMismatch becomes one that ends with the call to run again. A CSV the inference ladder cannot parse keeps polars' first paragraph (the value and the column), drops the advice after it, and ends with the suggested schema as a catalog_import_source call with schema=[[...]]. io.py is unchanged: source_import reads the ladder's message. The MCP tool catches every exception, so each failure gets an error_id, an errors.jsonl record with its traceback, a build_error event and a build_failed notification. test_suggested_schema_recovery_is_pasteable_with_reserved_column now reads the suggestion out of that call. #234: the append-only refusal names `tallyman revisions` and `tallyman reset-to <step>`, run from a shell. #239: the existing-entry branch restores a missing clone whenever it is missing, verified against the digest, at the path the entry names; ensure_cas_path takes the suffix, since the same bytes can arrive as .pq for .parquet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…he file is named #224, reworked: a parquet column xorq has no ibis type for is left out of the snapshot and named (in the return value, the recipe header and the recorded reader), not refused. The type check reads the file it names even when the name holds glob characters, and checks the clone, not the file before it was digested. Also, from the review of #247: a bug in the snapshot writer keeps its type, a failed alias write leaves nothing behind, and every failure of the MCP tool (a schema it cannot convert, a step after the head advanced) is recorded. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…alog #234 made the refusal name `tallyman reset-to <step>`, a command that exists, but it moves every alias and entry back to put one source on an older version. Until one alias can be moved back on its own, the refusal offers only the pinned read. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#224, reworked. A parquet column with a field xorq has no ibis type for (fixed_size_binary, so a UUID) is left out of the snapshot instead of refusing the file. The reader records it as `omitted`, so the entry hash covers it and the heal leaves it out too. The recipe header names each one, and the import returns `omitted_columns` and a `warning` with a cast that works: binary, or each UUID formatted with str(uuid.UUID(bytes=value)), since arrow refuses the cast of raw UUID bytes to string. A file with no other column is still refused. The check gave the caller's path to DataFusion's register_parquet, which takes a glob: `ids[v2].parquet` matched nothing and passed, and `ids[12].parquet` read the sibling `ids1.parquet`. It now reads a link with a plain name. The clone is checked again, so a file rewritten between the check and the digest is refused. From the review of #247: - a snapshot writer's own bug (a KeyError, say) keeps its type instead of becoming "could not read your file" (#227); - catalog_import_source records a schema it cannot convert, and a failure after the head advanced is recorded as `after_import_error` while the advance is still committed (#227); - the append-only refusal no longer advises resetting the whole catalog to move one alias back; it offers the pinned read (#234); - recon_cas_path, which nothing called and which named the clone by the live file's suffix, is gone, with LostSourceVersion. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review of b8acc06, and what was done about each finding
Also changed, at Paddy's direction: the #234 refusal no longer advises CI: red runs 36041713861 (10 failed, 984 passed) and 36043124024 (11 failed, 983 passed), in both cases exactly the new tests. Fix run 36043866014: ruff, the fast suite (993 passed) and the integration suite (7 passed) all pass. 🤖 Generated with Claude Code |
…t-hardening Resolves the conflict in src/tallyman_mcp/server.py with #241: keeps this branch's _record_import_failure and takes #241's _entry_url, which resolves the companion's URL at call time and gives None when no server holds the data dir. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The guard counted every name, attribute and parameter under src/ as defined, so a message naming a tool that does not exist passed whenever a variable was spelled the same way. Its scan moves into _unknown_catalog_names so a test can run it over a small tree. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…failed build The steps after catalog_import_source's import shared one try, so carrying the entry's config forward failing skipped the new_entry notification and the recalc, and the source's dependents stayed on the old version's rows. The failure was recorded as a build_error with no hash, so the Log showed a failed build for an import that was committed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ot a failed build The notebook append, the config carry-forward and the recalc each run in their own try, so one failing no longer skips the new_entry notification or the recalc. What failed is recorded once, as an after_import_error event rather than build_error, with a message that says the version was imported and with the new entry's hash, so a dependent left stale is tied back to it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…as defined A variable, parameter or attribute spelled like a tool is not something a message can send an agent to. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review of #247 (2026-09-28)A /code-review pass found ten things. Each was reproduced by a test that CI saw fail first ( Fixed here:
Split out:
🤖 Generated with Claude Code |
Stacked on #189: the base is
feat/adr-007-009-cache-redesign, notmain. Five issues in the ADR-011 import path:source_import.update_and_depend, which thecatalog_import_sourceMCP tool calls.Fixes #224
Fixes #225
Fixes #227
Fixes #234
Fixes #239
These close the issues only once #189 reaches
main.Terms: an import copies the user's file (the outside file) to a clone,
data/.cas/<md5><suffix>, then writes the snapshot, the parquet file of its rows atcompute_cache/result_cache/<hash>.parquet, then the entry directory with a generated recipe and a manifest. The recipe reads the snapshot withdeferred_read_parquet. While a source version's clone is gone, its snapshot is pinned: the Cache page refuses to delete it, since nothing else can make it again.What changed
#224: a column xorq has no type for is left out of the entry, and named
deferred_read_parquetasks xorq's DataFusion backend for the file's arrow schema and converts each field withPyArrowType.to_ibis, which has no entry forfixed_size_binary. Afixed_size_binaryor UUID column failed there with aKeyError, after the clone and the snapshot were written.update_and_dependnow asks that backend for the file's schema before it digests anything, and runs the same conversion field by field. A top-level column with any field that fails is left out of the snapshot, not refused.metagoes whenmeta.keyis afixed_size_binary. The import says so in three places:omitted,[[column, why]]. The entry hash covers it, and a heal from the clone replays it, so the healed snapshot's digest matches.# not imported:line per column.omitted_columnsand awarning:The advice no longer says to cast a UUID to string: arrow refuses that cast (
ArrowInvalid: Invalid UTF8 payload), since raw UUID bytes are not UTF-8. A file with no column left is still refused, and nothing is written.These messages can be missed. A note that stays with the source and is shown in the UI is #254.
The check reads the file it names.
register_parquettakes a glob pattern, not a path. Soids[v2].parquetmatched nothing, the schema came back empty, and the check passed afixed_size_binarycolumn.ids[12].parquetmatched the siblingids1.parquetand checked that file instead. The backend is now given a symlink to the file, with a plain name, in a temporary directory. The same glob reading makes every entry read as empty when the project path holds[,*or?. That is #255, not fixed here.The clone is checked again. The check reads the caller's file before it is digested. A file rewritten in between was checked on its old bytes and minted from its new ones.
_mintnow takes the schema of the verified clone. If its left-out columns differ from those recorded, the import is refused and the call to run again is printed.Deviation: the check reads DataFusion's schema, not
pq.read_schema. The issue suggested convertingpq.read_schema(src). The two schemas differ on extension types. DataFusion gives a top-level extension column its storage type and keeps a nested one:pq.read_schematypeto_ibisof itextension<arrow.json>string, importsextension<arrow.bool8>int8, importsextension<...>array<float64>, importsextension<arrow.uuid>fixed_size_binary[16], KeyErrorstruct<j: extension<arrow.json>>A check over
pq.read_schemawould leave out the first three, which import today. So the check asks the backend the read asks, through the same callxorq_datafusion.Backend.read_parquetmakes.test_the_type_check_passes_every_type_the_read_takescovers the first three rows and every other type the issue lists as importing. The check executes nothing, so it adds no execution site for #118.#225: a failed import removes what it wrote, and only that
_mintnotes which of the clone and the snapshot were absent when it started. One failure handler removes those, plus the entry directory if_mintcreated it. The project lock, whichupdate_and_dependholds, keeps anything else from writing those paths meanwhile. Two things stay:The repair branch, for an entry that already exists, restores a missing clone (#239) and keeps it even if the heal after it fails. That clone is named by a live entry, so it is not an orphan.
#227: every failure is recorded, and the message is about the user's call
SourceImportError: polars'ValueError,TypeError,NoDataErrorandComputeError, parsy'sParseError, and the schema DSL's own errors. It names the outside file andcatalog_import_source, not the clone ortallyman_read_csv. Only those types are wrapped (_reader_errors), and aTypeErroronly when the call passed reader options. Any other exception, a bug in tallyman's own writer, keeps its type, so the tool reports it with its class name.CloneDigestMismatch, from ADR-011 D9 (ingest verifies what it wrote), becomes one too. It ends with the call to run again.error_id, anerrors.jsonlrecord with the traceback, abuild_errorevent and abuild_failednotification, as_run_and_recorddoes for builds. An exception that is not one of the three expected types keeps its class name in the message. The conversion of a listschemais inside the handler too, andschema=[1]gets a message that says what a schema is.alias_setevent, the notebook, carrying the entry's config forward, the recalc) are handled as well. The head has already advanced by then, so the import happened: the reply is the import's, withafter_import_error: {error, error_id}recorded as above. The dispatch checkpoint still commits the advance. Before, a failure there skipped the checkpoint and left the head advance uncommitted.A CSV the inference ladder cannot parse, before and after:
The retry comes from
import_call, so it carriesreader_optionsandpinned_versionwhen the call had them. The test runs it as written, through the client, and it imports.Deviation: all of polars' advice is dropped, not only the lines about refused options. The message keeps polars' first paragraph, which names the value and the column. The advice after it is written as
scan_csvkeywords, not the tool's arguments. Besides the two options the import refuses, it suggestsignore_errors=True, which would turn the bad values into nulls without saying so, andnull_values, which the caller can pass inreader_options.io.pyis unchanged, to stay clear of the_polars_dtypeand timestamp work on #231.source_importreads the ladder's message with a regex (_LADDER_FAILURE) and strips thetallyman_read_csv:prefix from the others. The cost is that the schema DSL's own messages still show their examples as tuples, as inschema=(("Date", "date"), ("&rest", "infer")). Those are examples of the positional form, not a retry call.#234: the append-only refusal does not send the user to a command that does not exist
It first said to reset with
tallyman resetorcatalog_reset_to, and neither exists. It no longer advises a reset at all.tallyman reset-to <step>exists, but it moves the whole catalog back, every alias and entry with it, to put one source on an older version. The refusal now offers onlypinned_expr_from_alias('<alias>-v<N>'), which reads the old version without moving the head. Moving one alias back on its own is #256, which also amends ADR-011 D11.test_every_catalog_name_a_message_in_src_quotes_is_a_tool_or_a_python_nameis the guard the issue suggested. It covers string literals insrc/, but not docstrings, sincedependents.py's docstring names the deletedcatalog_parentsandcatalog_dagas history. It does not coverREADME.md, whosecatalog_load_parquetmentions #216 rewrites.#239: a re-import restores a lost clone whenever it is missing
The existing-entry branch calls
ensure_cas_pathevery time. That function writes the clone only when it is missing and verifies it against the digest (D9). The clone goes to the path the entry recorded:ensure_cas_pathnow takes the suffix. The same bytes imported fromorders.pqfor an entry minted fromorders.parquethave the same entry hash. Before, a repair from such a file wrote a second clone under.pqand left the entry's clone missing.Tests
In
tests/test_source_import.py, a new section at the end, and one guard intests/test_mcp_tool.py:fixed_size_binary, a struct-nested one and a UUID column imports with onlyn. It returnsomitted_columnsand awarningnaming all three, with a cast that works. The recipe header and the recorded reader name them, the alias reads throughtracked_expr_from_alias, and a deleted snapshot heals with the same columns. The same, for files namedids[v2].parquet,ids[12].parquet(beside a decoyids1.parquet),ids*.parquetandids?.parquet. A file rewritten between the check and the digest is refused and leaves nothing.fastmcp.Client(in memory), seven CSV failures each return{error, error_id}, with oneerrors.jsonlrecord, onebuild_errorevent and abuild_failednotification. The seven are a pinned-type mismatch, a ragged row, invalid UTF-8, a dtype that does not parse,separator="ab", an unknown reader option, and an empty file. The message names the file and the tool, nottallyman_read_csvor the clone. For the mismatch, the advised retry is parsed and run as written. A clone write that fails its digest check is recorded the same way and leaves no clone.schema=[1]is recorded the same way. A failure after the head advanced is recorded asafter_import_error, and the step still moves. AKeyErrorfrom the snapshot writer keeps its type and leaves nothing.catalog_*guard above.pinned_reasonreturnsNone, from a file with the same suffix and from one with another.test_suggested_schema_recovery_is_pasteable_with_reserved_columnnow takes the suggestion out of the advised call, since the message no longer ends inschema=<tuple>.test_recon_cas_path_raises_instead_of_serving_drifted_live_bytesis deleted withrecon_cas_path, which nothing insrc/called, and which named the clone by the live file's suffix (the bug #239 fixed for the repair path).How it was built
b284b6dadds the 16 failing tests. CI run 36017772166: ruff passed, and the fast suite had 16 failed and 954 passed. The 16 failures were exactly the new tests, each for the reason its issue gives: aBuildErrorcarrying theKeyError(importing a parquet file with a fixed_size_binary column fails with a KeyError traceback that names no column #224), leftover files (a failed import leaves its clone in data/.cas and its snapshot in result_cache/ — nothing ever deletes the clone #225),is_errorwith no record (catalog_import_source lets a CSV reader error escape unrecorded — the message names the clone and tallyman_read_csv #227),catalog_reset_toin the text (the append-only import error names reset commands that do not exist (tallyman reset, catalog_reset_to) #234), a clone still missing (re-importing a version whose clone is lost does not restore the clone while its snapshot exists, so the version stays pinned #239).e6447b9fixes importing a parquet file with a fixed_size_binary column fails with a KeyError traceback that names no column #224.b8acc06fixes the other four. CI run 36020422748 on the tip: ruff, the fast suite (972 passed, and vitest 15 passed) and the integration suite (7 passed) all pass.667557dadds 10 failing tests and026148fchanges the the append-only import error names reset commands that do not exist (tallyman reset, catalog_reset_to) #234 test. CI runs 36041713861 (10 failed, 984 passed, exactly the new tests) and 36043124024 (11 failed, 983 passed).e0e91b9reworks importing a parquet file with a fixed_size_binary column fails with a KeyError traceback that names no column #224 and fixes the rest. CI run 36043866014: ruff, the fast suite (993 passed, and vitest 15 passed) and the integration suite (7 passed) all pass.Checked locally before each push, with
TALLYMAN_HOMEandTALLYMAN_COMPANION_URLpointed at scratch values:test_fouc::test_unknown_api_path_404sfailed because it needs the built React app, as on fix(xorq): a failed build keeps the snapshot already on disk for its hash (ADR-011 port of #213) #222 and fix(cache): a pin survives a reset and a dismissed error banner (#194, #195, #196), on the ADR-011 stack #223.uvx ruff checkpassed.The repo has no
.pre-commit-config.yaml, and the files this PR edits are not ruff-formatted as a whole. Soruff formatwas applied to the lines this PR touches and nowhere else.Not in this PR
a source imported without some columns carries no permanent note saying so #254: a note on the source, shown in the UI, that some columns were not imported.
every entry reads as empty, with no error, when the project path holds a glob character #255: with
[,*or?in the project path, every entry reads as empty, with no error.no way to move one source alias back to an older version #256: moving one source alias back to an older version.
An alias write that fails after
_mintreturns leaves the new entry, snapshot and clone on disk with no alias. Found in review, and deferred.an unfaithful heal of a source version is blamed on the recipe (#83), and a corrupt clone is healed from and pinned #232: a corrupt clone is still trusted, since
ensure_cas_pathwrites only a missing one.a writer killed mid-write leaves a .tmp file in result_cache/ that the Cache page never lists and nothing deletes #226: temp files a killed writer leaves in
data/.cas/andresult_cache/.For a ragged row, the suggested schema does not fix the file. That was already so, and tallyman_read_csv: ragged/short CSV rows are silently null-filled instead of raising #142 covers the short-row side.
A failed import can write a clone that a different entry has lost (the same bytes, read with other reader options) and then remove it, since it was absent when the import started. That entry stays pinned until its own re-import restores the clone (re-importing a version whose clone is lost does not restore the clone while its snapshot exists, so the version stays pinned #239).
🤖 Generated with Claude Code