Repository navigation
Conversation
Proposes two changes so a heal is flagged only when the result changed: materialization runs single-partition, because a float SUM/AVG is not bit-reproducible under DataFusion's parallel partial aggregation, and result_digest becomes a digest of the snapshot's ordered Arrow content, computed from the file as read back, because file bytes move with the writer's version and with how the stream was batched. Amends ADR-004 and ADR-006 D5. scripts/spike_float_aggregate_digest.py: 6 distinct digests in 6 runs at default partitions, 1 at target_partitions=1, integer control stable. scripts/spike_logical_digest.py: the content digest is invariant to batch size, codec and row-group size, and detects a changed value, a null replaced by 0.0, and a row swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This was referenced Sep 18, 2026
…nts from ADR-008 Adds a Terms section, writes every other ADR's decision label with its ADR number and what it decides, and replaces "bake" with "materialize". D3 (the snapshot's format) gains two requirements from the revised ADR-008: the writer numbers rows in a last column named __row_order, and it writes a parquet page index, which takes a range request from 79-90 ms to 19-24 ms at any depth. The digest covers __row_order like any other column. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Paddy, in the grilling session: a cohesive system that works reliably comes first, and speed problems are handled as they come up. Every materialization runs single-partition. The float-only variant is noted as a later option and is no longer a gate on the decision. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Draft ADR for discussion. Adds
plans/ADR-009-digest-stability.mdand the two spikes it cites,scripts/spike_float_aggregate_digest.pyandscripts/spike_logical_digest.py. No behaviour change.A heal whose digest does not match is loud and sticky: a durable
unfaithful_healrecord, an SSE event, a stat-cache wipe, session eviction, and a pin and badge on the entry (ADR-006 D7, D10, D12). Two things trigger that today without any change in the result.SUM/AVGis not bit-reproducible under DataFusion's parallel partial aggregation. The same canonically ordered group-by over 3,000,000 rows gives 6 distinct digests in 6 runs at the default partition count and 1 attarget_partitions = 1; an integer-only aggregate is stable either way. Aggregate is the usual reason an entry is worthy, so most expensive entries with a float measure are flagged after any eviction.Decisions proposed:
target_partitions = 1, for the build and every heal. About 3x on aggregation at spike scale, paid per bake and never per read. Gated on measuring a parking-corpus rebuild, with a recorded float-only fallback.result_digestbecomes a SHA-256 over the snapshot's ordered Arrow content, computed from the file as read back by one function shared by build and verify. It is invariant to batch size, codec and row-group size, survives a trip through DataFusion, and still detects a changed value, a null replaced by0.0, and a row swap. It costs 0.12 s to read back and hash a 57 MB file against 0.02 s for a file-bytes hash. This reverses part of ADR-006 D5's reasoning, and the ADR says so.__row_order, and it writes a parquet page index, which takes a range request from 79-90 ms to 19-24 ms.unfaithful_healrecord carries engine and writer versions at build and at heal, so an upgrade is not blamed on the recipe.One of three ADR PRs from the 2026-09-18 cache audit, with ADR-007 (#180) and ADR-008 (#181). References to the other two do not resolve until those merge. D2 and D3 assume ADR-007's tallyman-side writer; D1 stands without it.
Checked locally:
ruff checkandruff format --checkon both scripts, both run from a clean temp dir and their output matches the tables in the ADR, and every path and line reference in the ADR resolves on a tree with all three ADRs present. Docs-only, so CI was not watched.🤖 Generated with Claude Code