From af38dced9d2e8188768407ef569cc5d25549417e Mon Sep 17 00:00:00 2001 From: Scout Date: Sat, 26 Sep 2026 14:15:08 -0400 Subject: [PATCH 01/15] Mover v3 round 18: pre-freeze residual coverage, seal recovery for dry runs and live samples, literature correction Residuals the round-17 verifiers and the GPT-6 re-review named: - R1-R4, coverage only. Eight tests pass on unchanged 803bc351 production code: - the H3-c split on the last overnight leg, and a split and a spin-off on different legs; - a name_change dated exactly on asof t, for every issuer-record consumer and the terminal-zero rename diagnostic; - the exact terminal_actions refusal reason; - a direct test of the no-valid-page-calibration guard. - R5: count_only.dry_run and transport_check.seal_live_samples adopt the recover-if-matching pattern of core.holdout.collect. An existing snapshot is validated with Store.read; its binding and complete request/stamp set must match the new request; its sha256 is reused without fetching or writing. Partial, mismatched, unbound or altered snapshots refuse. Eight tests are in tests/test_seal_recovery.py. Reverting only the two production files gives 30 failures and 1 error. The dry-run local variable is named seal_identity so it does not shadow the imported core.identity module. - Literature: item_set.items[4].why_kept described the LPS 2019 and ABJK 2022 samples together as 'large, liquid names'. Re-verified from the manuscripts: LPS excludes microcaps (below $5 or the bottom NYSE size quintile) and opens at the first-half-hour VWAP; ABJK keeps common stocks above $1 (excluding financials and utilities), with its effect in all but the largest decile. The test stays two-sided; no rule changes. Two round-18 review_record entries. Study test_command: 376 tests OK. The protocol stays draft_pending_independent_pre_outcome_review, with frozen_before_outcomes false. R1-R5 were implemented by GPT-6 (gpt-6-astra, effort max) test-first and verified by the coordinator; the literature correction was found by the GPT-6 paper sweep and verified by the coordinator. Co-Authored-By: Claude Opus 5.5 (1M context) --- .../mover-v3/protocol-core-draft.json | 18 +- .../mover-v3/study/core/count_only.py | 29 ++- .../mover-v3/study/core/transport_check.py | 28 ++- .../mover-v3/study/tests/test_calibration.py | 9 + .../mover-v3/study/tests/test_holdout.py | 22 ++- .../study/tests/test_seal_recovery.py | 175 ++++++++++++++++++ .../mover-v3/study/tests/test_trades.py | 68 +++++++ manifests/evidence.json | 4 +- 8 files changed, 327 insertions(+), 26 deletions(-) create mode 100644 blueprints/us-equities/mover-v3/study/tests/test_seal_recovery.py diff --git a/blueprints/us-equities/mover-v3/protocol-core-draft.json b/blueprints/us-equities/mover-v3/protocol-core-draft.json index fad8b6f95..c59d2ab44 100644 --- a/blueprints/us-equities/mover-v3/protocol-core-draft.json +++ b/blueprints/us-equities/mover-v3/protocol-core-draft.json @@ -62,7 +62,7 @@ "role": "primary test (not a tradable cell)", "statistic": "mean over events of (mean of the 4 overnight leg log returns - mean of the 4 intraday leg log returns), gross, legs as in arms.c_decomposition", "alternative": "not equal to 0 at every stage; at the holdout a pass also needs the sign of the validation estimate (multiple_testing.no_direction_lock)", - "why_kept": "The direct test of H3's 'versus'. It uses official prints only, so it needs no fill and no cost model. The literature (LPS 2019, ABJK 2022) concerns large, liquid names and gives no sign for movers, so the test is two-sided at every stage. It is not locked to the development sign, because development is survivorship-limited (chronology.labels.development) and gates nothing. At the holdout, a confirmatory pass also needs the holdout estimate to have the sign of the governing validation estimate, so a pass replicates the screened direction." + "why_kept": "The direct test of H3's 'versus'. It uses official prints only, so it needs no fill and no cost model; it is therefore not an exact replication of LPS 2019, whose open is the first-half-hour VWAP. LPS 2019 excludes microcaps (price below $5 or the bottom NYSE size quintile), and ABJK 2022 keeps common stocks above $1 (excluding financials and utilities) with its effect in all but the largest size decile; neither studies the mover population or gives a sign for movers, so the test is two-sided at every stage. It is not locked to the development sign, because development is survivorship-limited (chronology.labels.development) and gates nothing. At the holdout, a confirmatory pass also needs the holdout estimate to have the sign of the governing validation estimate, so a pass replicates the screened direction." } ], "tradable_cells": "H1-D-b_lane-low, H3-a and H3-b are tradable cells and nothing else. H1-D and H3-c are primary tests and are never tradable cells. The robustness requirement (statistics.robustness_required_for_a_pass) applies to the three tradable cells only.", @@ -862,7 +862,7 @@ "fetch": "Before any outcome of a stage (development, validation, and each holdout 'count' or 'read'), the frozen code fetches every quote window, minute bar, per-event daily bar and official print that stage needs, and for development and validation first the full 2016-2020 screen data (universe_and_identity.request_shapes), under universe_and_identity.asof. Review round 8, R8-1: the request plan (symbols, asof, sessions, stamps, entry, exit, search and backward windows, and the enumeration) is derived by core/plan.py and core/stage.py, in the evaluation tree, only from sealed inputs (the calendar, the enumeration, corporate-action records, decision-time inputs, and the timestamps and eligibility of already sealed quotes); the loop core/driver.py fetches every planned request not yet sealed until the plan adds nothing. study/fetch/ receives one planned request and returns its raw pages; it derives nothing and parses no row. Terminal-search windows of censored trades extend past the segment end (populations.terminal_exits.segment_end), and backward windows follow populations.terminal_exits.booking (R8-7). The result is sealed as a sha256-pinned private snapshot (the ledger's sha256; the ledger holds every page's sha256 and normalized-record sha256), and that sha256 and the stage's fetch-incomplete rate by kind are committed in the run log. At most one re-fetch is allowed: once, of only the fetch-incomplete requests, before any outcome of that stage, followed by fetching any window that only then became plannable, committed the same way. After it, the 1% void rule is final. No outcome is computed from unsealed data.", "synthetic_tests": "The evaluator refuses a results file whose run-log entry names a tree other than the pinned one, a holdout count or read from a tree other than the pinned one or a passing transport deviation of it, a transport deviation without a passing reproduction check, a runtime that differs from runtime.lock, an edited (not appended) amendment file, a second results file for a stage, a read retry that chronology.holdout.read_timing does not permit, and an outcome computed before the stage's fetch rate is committed. It accepts a run from a commit reachable from origin/main whose study tree equals the pinned tree, with run-log, access-log and amendment lines committed after the freeze. A test kills a run before its results write and asserts that nothing was written and that the retry reproduces an uninterrupted run byte for byte. The suite is study/tests/, run by study_code.test_command under study/runtime.lock. Review round 9 (F6) adds command-level tests against a temporary repository shaped like this one after the freeze (tests/fixture_repo.py), so no test depends on the real protocol's status or writes to the real run log: every command refusing on a draft, a stage fetched once and evaluated into one results file with a kill after the write, the context refusals (bytecode isolation and ignored files, logs off origin/main, a protocol edited after the freeze, amendment timing, coverage_decision, transport deviations), the count-only command, and the holdout path through collect, count with two extensions and read, with two batches of different enumerations and a rename re-fetch, and the not-read label after the deadline.", "pinned_copies": "The study tree imports no module outside itself except the Python standard library, numpy, pandas, exchange_calendars and duckdb (review round 7, F6; review round 8, E1 adds duckdb, which the pinned dedupe_identity needs). It holds byte-for-byte copies of the external definitions it needs, listed in study/pinned_copies.json as (source path, commit aa6fc79, git blob, sha256, definition names): the Bars class and ET, et_epoch, TIME_BUCKETS, PRICE_TIERS, DV_TIERS, time_bucket, price_tier and dv_tier from mover-early-entry rules.py (R8-12); official_price from sessions_io.py; valid_quote from quotes.py; dedupe_identity with DUPLICATE_SERIES_MIN_SESSIONS, DUPLICATE_SERIES_MIN_SPAN_COVERAGE, CONTINUITY_RATIO, _rename_terminus and _components from broad-universe coverage.py; DATA_SYMBOL from collect_daily.py. A synthetic test reads each source with git cat-file at the listed blob, checks the blob hash and sha256, extracts each named definition with ast (decorators included) and asserts that the copy is identical and holds nothing else. Because nothing outside the tree is imported, a later edit to those paths on main cannot reach a run, and no module-level code of rules.py, quotes.py or coverage.py (their fees.json, protocol.json and duckdb-connection code) ever runs.", - "transport_deviations": "The fetch transport must work against a live API for 12-18 months (review round 7, C2). Review round 8, R8-1: study/fetch/ holds only the transport (HTTP client, host, endpoint path, authentication header, pagination and retry); the request plan, the stamps and windows, the enumeration, response parsing and row filtering are evaluation code outside it, so no change under study/fetch/ can change a request, a stamp, a window, the parsing or the evaluation. It can change which bytes arrive (review round 11, F3: the child process returns raw pages that the parent parses), which the reproduction check below bounds by sampling and does not exclude. A deviation that changes only files under study/fetch/ is a transport deviation. Before any run from the new tree, it passes a committed reproduction check against the most recently sealed snapshot of each request kind that exists (or the snapshot of the pre-freeze native dry run, freeze_preconditions): (a) the new tree regenerates the request plan from the sealed inputs and its canonical request records equal the sealed request records byte for byte; (b) every sealed raw page, re-parsed offline, gives byte-identical normalized records (its sealed normalized_sha256), for every page, not a sample; (c) the new transport re-requests the first 200 requests of each kind, in sha256 order of their request records, and their merged normalized records equal the sealed ones (core/transport_check.py). The result is committed in deviations.json with the deviation. A passing transport deviation may govern every later fetch, collection batch, count and read, whose results carry the qualifier 'transport-deviation'. A deviation that touches any file outside study/fetch/, or fails the check, is an ordinary deviation (deviations). Limitation: a change of the provider's response format needs a parser change outside study/fetch/, which is an ordinary deviation. If no passing transport deviation restores collection or fetching in time, the affected sessions are late-collected (holdout_gate.prospective_collection) or the read misses its deadline, and chronology.holdout.read_timing applies. Record format (review round 9, H-2): deviations.json is {\"deviations\": [...]}, and a transport deviation carries number, kind 'transport', cause, diff_reference, new_tree, new_fetch_tree, affected_stages and transport_check (the reproduction check's result, with passes). core.guards.check_study_tree accepts a running tree other than the pinned one only when a committed entry names it as new_tree and the running study/fetch/ tree as new_fetch_tree, its transport_check passed, and the pinned tree and it differ only in paths under fetch/. Every result from it carries 'transport-deviation'. Review round 10, H1 and F6: the reproduction check runs only as run.py transport-check from the new tree (which may differ from the pinned tree only under fetch/), against a development or validation snapshot sealed by that stage's logged fetch (it holds every screen and per-event request kind). It writes results/transport-check-.json with a 'transport_check' run-log line. transport_check in the deviation must carry passes, output_path and output_sha256, and core.guards.check_study_tree accepts the deviation only when that committed output says passes for the new tree and a complete run-log line from the new tree wrote it; a self-declared passes: true is refused. Because the transport runs in a child process that returns raw bytes only, a change under fetch/ cannot alter planning, parsing, seeds, statistics or labels, and (a) and (b) need no rerun at run time: they are deterministic functions of unchanged code and sealed bytes, bound by the output's hash. Limitation: (c) keeps the fixed sample of the first 200 requests per kind in sha256 order of their request records. Any committed rule is known in advance, and an unpredictable sample could not be reproduced, so a transport written to alter only unsampled responses is not caught by (c). A stage snapshot holds no enumeration request (assets and corporate actions come from the count-only part 0 and the collection batches), so transport-check does not re-request those kinds, and a change that alters only them is not caught by (c). Review round 11, F3: (c) samples 200 requests per kind in sha256 order of the seed followed by the request record, where the seed is the id of the first origin/main commit holding the new tree, a signed merge made after the transport was written, so the sample cannot be known when the transport is written. transport-check also runs (b) and (c) against every sealed holdout snapshot (collection batches and count and read snapshots) and refuses without --holdout-root once one exists; (a) is not rerun for them, since it is a deterministic function of the unchanged plan code. Development is evaluated only after validation's fetch line is on origin/main, which closes the window between development results and the validation fetch. A holdout read under a passing transport deviation may support a claim, with the qualifier 'transport-deviation'. Limitation: a transport that alters responses of unsampled requests only is still not caught, now with a probability set by the unpredictable sample. Review round 15, N02 (R14-open-2): run.py transport-check seals each live sample (c) under /transport-check--live/