Skip to content

Mover v3 round 18: pre-freeze residual coverage, seal recovery for dry runs and live samples, ABJK/LPS literature correction — not frozen - #360

Merged
seathatflowsinourveins merged 16 commits into
mainfrom
claude/mover-v3-round18-20260926
Oct 1, 2026
Merged

seathatflowsinourveins merged 16 commits into
mainfrom
claude/mover-v3-round18-20260926

Conversation

@seathatflowsinourveins

Copy link
Copy Markdown
Owner

Status (2026-09-26): pre-freeze residuals, not frozen; independent reviews in progress

Round 17 (#351, GPT-6 ready_to_freeze) left residuals: coverage the verifiers named, and one fail-closed usability gap. This PR closes them and makes one literature correction. The protocol stays draft_pending_independent_pre_outcome_review with frozen_before_outcomes: false. No outcome data, sealed file or market data was opened.

Round 18

Item Kind Change Evidence
R1 (F1) coverage H3-c split on the last overnight leg; split and spin-off on different legs pass on unchanged 803bc35
R2 (F2) coverage a name_change exactly at asof t, for every Ctx.issuer_records consumer and the terminal-zero rename diagnostic pass on 803bc35
R3 (F3) coverage the exact terminal_actions refusal reason and the planner boundary pass on 803bc35
R4 (F4) coverage a direct test of the no-valid-page-calibration guard pass on 803bc35
R5 (F5) production count_only.dry_run and transport_check.seal_live_samples recover a matching sealed snapshot and refuse partial, mismatched, unbound or altered ones, with no refetch or overwrite (the holdout.collect pattern) 8 tests in test_seal_recovery.py; reverting the two files gives 30 failures and 1 error
Literature text only item_set.items[4].why_kept described the LPS 2019 and ABJK 2022 samples together as "large, liquid names" re-verified from the manuscripts; the test stays two-sided
  • LPS 2019 (LSE accepted version): microcaps are always excluded, meaning a price below $5 or the bottom NYSE size quintile. The open is the first-half-hour VWAP.
  • ABJK 2022 (JFE 145(3), September 2021 manuscript): CRSP share codes 10 and 11, excluding financials, utilities and stocks at or below $1, May 1993 to December 2017. The effect holds in all but the largest size decile.
  • Sweep: a live GPT-6 academic-paper sweep for strategy-research found it (25 verified papers; evidence/artifacts/trading-convergence-20260926/papers-strategy.md, landing in Trading research convergence 2026-09-26: 13 layers, GPT-6 live refutation, verbatim-checked votes (verdict-wave input) #358).

The two new round-18 review_record entries cover R1–R5 and the literature correction.

Checks

  • Study test_command: 376 tests OK (359 at 803bc35 plus the new tests).
  • R5 revert check: 30 failures and 1 error across the 4 seal-recovery test methods.
  • scripts/validate.py passed. The only evidence change is the protocol pin; study code and tests are not hash-listed.
  • Authorship: GPT-6 (gpt-6-astra, effort max) implemented R1–R5 test-first. The coordinator verified them and renamed one local variable in dry_run that shadowed the imported core.identity module. The coordinator made the literature correction.
  • Reviews in progress: an independent read-only Claude review (Opus) of the whole delta, and a GPT-6 read-only live-search delta review as the pre-outcome reviewer of record.

Path to a freeze (unchanged)

  1. F12 budget values, set after the native dry run by a reviewed PR. The dry run needs Alpaca market data; paper keys currently return 401.
  2. The freeze commit, once freeze_preconditions hold.

SOTA sources

  • ABJK 2022: doi:10.1016/j.jfineco.2021.09.019 (manuscript: Akbas, Boehmer, Jiang and Koch, September 2021)
  • LPS 2019, A tug of war: https://eprints.lse.ac.uk/87481/
  • core.holdout.collect recover-if-matching at 803bc35 (this repository), the pattern R5 follows

🤖 Generated with Claude Code

…y runs and live samples, literature correction

Residuals the round-17 verifiers and the GPT-6 re-review named:
- R1-R4, coverage only. Eight tests pass on unchanged 803bc35 production code:
  - the H3-c split on the last overnight leg, and a split and a spin-off on
    different legs;
  - a name_change dated exactly on asof t, for every issuer-record consumer
    and the terminal-zero rename diagnostic;
  - the exact terminal_actions refusal reason;
  - a direct test of the no-valid-page-calibration guard.
- R5: count_only.dry_run and transport_check.seal_live_samples adopt the
  recover-if-matching pattern of core.holdout.collect. An existing snapshot
  is validated with Store.read; its binding and complete request/stamp set
  must match the new request; its sha256 is reused without fetching or
  writing. Partial, mismatched, unbound or altered snapshots refuse. Eight
  tests are in tests/test_seal_recovery.py. Reverting only the two production
  files gives 30 failures and 1 error. The dry-run local variable is named
  seal_identity so it does not shadow the imported core.identity module.
- Literature: item_set.items[4].why_kept described the LPS 2019 and ABJK 2022
  samples together as 'large, liquid names'. Re-verified from the
  manuscripts: LPS excludes microcaps (below $5 or the bottom NYSE size
  quintile) and opens at the first-half-hour VWAP; ABJK keeps common stocks
  above $1 (excluding financials and utilities), with its effect in all
  but the largest decile. The test stays two-sided; no rule changes.

Two round-18 review_record entries. Study test_command: 376 tests OK. The
protocol stays draft_pending_independent_pre_outcome_review, with
frozen_before_outcomes false. R1-R5 were implemented by GPT-6 (gpt-6-astra,
effort max) test-first and verified by the coordinator; the literature
correction was found by the GPT-6 paper sweep and verified by the coordinator.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins seathatflowsinourveins added the lane:trading Trading lane: sim, paper and live north star label Sep 26, 2026
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

GPT-6 cross-family pre-outcome review of af38dce (gpt-6-astra, effort max, read-only, live search): VERDICT: ready_to_freeze.

  1. R5 is safe. Recovery validates the seal, the exact binding, the canonical requests and the request/stamp coverage. Matching recovery performs no fetch or write, and mismatches refuse.
  2. The literature correction is accurate and text-only. Live searches confirm LPS's microcap exclusions and 9:30–10:00 VWAP open. They also confirm ABJK's rules: common shares with a previous month-end price above $1, financials and utilities excluded, and significant hedge returns outside the largest size decile. No statistical rule, arm, horizon or threshold changed.
  3. No new freeze blocker. Its sandbox couldn't run the pinned suite (uv cache lock), so the 376-test result remains the coordinator's evidence.

The independent Claude review is still running.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

The independent Claude review of round 18 was interrupted by a Claude session limit (HTTP 429, resets 22:10Z). R1–R5 were written by GPT-6, so the cross-family review must come from Claude; no other reviewer substitutes. This PR waits for that review before merging. Nothing is time-critical: the freeze is blocked on the market-data dry run in any case.

seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
…t-span quotes, survival rule wording, round-18 status, privacy)

GPT-6 read-only pre-merge check of 260963f (needs_changes), addressed:
- Quote integrity: 223 of 716 stored quotes had typographic punctuation
  normalized. Quotes are now the source's own span (located tolerating only
  whitespace and typographic punctuation), sanitized exactly as the exported
  texts, and kept only if they occur in the published text. 716 of 716 are
  exact spans; none dropped.
- Survival rule: only position 'refutes' changes survival; 'corrects' fixes
  a fact or claim and leaves the label standing. The wording is now
  explicit.
- Round 18: the ABJK/LPS correction is described as proposed in #360, under
  review, not as done.
- Privacy: email addresses are replaced by <email>. Two claim checks that
  quoted private host memory keep their name and verdict with their text
  withheld.

The record's content and dispositions are otherwise unchanged
(95/40/47, 716 votes). validate.py, validate_catalogs.py and gitleaks are
clean; validate.py's private-content patterns find nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Scout and others added 5 commits September 30, 2026 20:28
Two command-level regressions for the dry-run seal recovery, through
run.py dry-run against the fixture repository. Each run's output step
fails after the seal, so run.py's finally block writes a failed end line
and the next run opens a new start line.
- H1: tree B, which changes only study/fetch/ and so keeps the same request
  plan, must not adopt tree A's seal. At af38dce it does:
  "AssertionError: SealError not raised" at test_seal_recovery.py:239.
- M1: the same tree retrying on the next UTC day must adopt its seal and
  read back the first run's start. At af38dce the fetch date is matched:
  "core.store.SealError: dry-run snapshot binding differs from the
  requested dry run" (count_only.py:486).
The binding case that asserted a fetch-date refusal (the wrong rule, the
review's M1) is dropped. Module run at af38dce: Ran 10 tests,
FAILED (failures=1, errors=1).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eal recovery

Resolves the round-18 pre-outcome reviews of the two recoverable seals
(count_only.dry_run, transport_check.seal_live_samples), following
core.holdout.collect's identity shape.

- G-H: Store.write(seal_record=True) publishes seal.json once, before the
  ledger. It holds the ledger sha256, the binding, every completion stamp
  [key, attempt, status, pages], every page sha256, and the attempt and
  page counts. core.store.recover_seal reads the ledger under the recorded
  digest, never a digest of the ledger being checked, and the snapshot
  must give the record back exactly. Store.read checks each stamp's page
  count against its page events.
- H1: each seal binds the running study tree, protocol sha256 and
  runtime-lock sha256 (runner.seal_run_identity, validated by
  store.checked_run_identity). The dry-run identity is {study_tree,
  protocol_sha256, runtime_lock_sha256, sessions, symbols, requests}; live
  samples add the same three fields to label, seed, source and requests.
- M1: the dry run's fetch start (utc_start) is bound beside the identity,
  read back into the output as fetch_utc_start, and never matched.
- G-M: separate checks for the identity, the request plan, and the ledger's
  request records against the requests the binding names.
- L1: attempt stamps are compared exactly through the seal record.
- L2: every recovery refusal is a SealError (zlib.error, TypeError,
  AttributeError and OSError included), naming the dry run or the label.
- L3: run-wide changes are sealed into a mixed live root, so each label's
  own refusal is asserted.

Tests: tests/test_seal_recovery.py, 15 new unit cases. The commit-1
command tests now pass. Callers in test_count_only.py and test_fetch.py
pass the run identity. Study suite: Ran 393 tests, OK.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…test record

- freeze_preconditions (native dry run) and run_discipline.transport_deviations
  now describe recovery through core.store.recover_seal:
  - verification against the seal record;
  - a binding to the running study tree, protocol sha256 and runtime-lock
    sha256 plus the caller's inputs and request plan;
  - the dry run's fetch start read back as fetch_utc_start, never matched;
  - a new snapshot root for any other tree, protocol or runtime lock.
- review_record: one round-18 entry dated 2026-10-01 for the Claude
  (H1, M1, L1-L5) and GPT (G-H, G-M) reviews. It records the failing-first
  output, the reproduction at af38dce, the mutants with their killing
  tests, the record's limit (a write-once file, not an external anchor),
  what is not changed (holdout.collect; orphan adoption), and the GPT
  review's confirmation of the LPS 2019 and ABJK 2022 sample definitions.
- L4: the round-18 literature entry cites PR #358 (not yet on main),
  7e29448:evidence/artifacts/trading-convergence-20260926/papers-strategy.md.
- L5: draft_tree_at_this_revision, test_record and freeze_acceptance[7]
  name tree 20a3dab (fetch tree
  unchanged) with 393 tests. test_record.command is the command actually
  run; command_note explains the environment prefix.

The study tree is unchanged by this commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…qualified recovery rules, review record

An independent verification of the round-18 repair passed it with five low
findings; four are repaired. The fourth, a stray git-ignored scripts/__pycache__
in the review worktree, was deleted and changes no tracked file.

- Protocol: the dry-run sentence of freeze_preconditions and
  run_discipline.transport_deviations no longer say that an altered snapshot
  refuses outright. A snapshot without its seal record is refused, and a
  partial, unbound, mismatched or altered one is refused when its seal.json is
  intact. seal.json is a write-once record inside the snapshot directory, not
  an external anchor: a consistent rewrite of both the ledger and seal.json
  that keeps the binding, the request plan and the requests is adopted.
- tests/test_seal_recovery.py: test_a_resealed_page_replacement_refuses (dry
  run, and each live-sample label) replaces a page with its ledger digests and
  reseals only ledger_sha256 in seal.json. It asserts 'the snapshot differs
  from its seal record', no transport call and unchanged bytes. The probe
  mutant that empties page_sha256s in core.store.seal_record_of passed all 393
  tests at 5a1d4b0; it now fails this test's three cases and nothing else
  (Ran 395 tests, FAILED (failures=3)).
- review_record, round-18 repair entry: one sentence names the three diff
  items it omitted (test_record.command with its command_note, Store.read's
  second-stamp refusal for every snapshot type, Store.write's refusal of a
  directory holding seal.json). Three residuals are recorded: the seal.json
  limit; core.holdout.collect reading its seal under a digest of its own ledger
  (predates this round); orphan-output adoption verifying seals against the
  sha256 values the output body names, not seal.json.
- run_discipline.study_code.draft_tree_at_this_revision, test_record and the
  freeze_acceptance evidence name the new study tree
  6cb88f0 and 395 tests (fetch tree
  f288cf2 unchanged).

Study suite: Ran 395 tests, OK.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
manifests/evidence.json is main's copy (3361b34) with protocol-core-draft.json
re-registered through scripts/host_receipts.register_file; component_matrix.py --write and
new_host_grand_list.py --write re-run (no change). scripts/validate.py and
scripts/evidence_manifest.py --check exit 0 on this tree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Independent pre-outcome reviews and repair (trading lane, 2026-10-01)

Head is now 406ad3c5: four repair commits on af38dced, then a merge of main at 3361b342 under the docs/lanes.md hot-file protocol (fast-forward push, no force). The protocol status stays draft_pending_independent_pre_outcome_review; nothing here reads an outcome.

Reviews of 803bc3514ef0..af38dced

Review Finding Severity
Claude Opus, effort max, read-only (2026-09-30) H1. The dry-run seal binding had no study tree, so a later tree could adopt an earlier tree's sealed snapshot without fetching. high
M1. fetch_date was part of the matched identity but came from the current run, so a rerun on a later UTC day was refused for good. medium
L1–L5. Attempt-stamp comparison; refusals that escaped SealError; a subTest that never reached its label; a citation of a file that is only on #358's branch; a stale tree and test count in run_discipline. low
GPT gpt-6.1-sol, effort ultra, native Codex 0.159.2, read-only (2026-10-01 00:03Z) G-H. Recovery trusted a digest it computed from the files it was checking. In the reviewer's synthetic run, a replaced page with matching ledger digests was adopted and turned passes=false into passes=true. high
G-M. Two recovery controls survived the tests as mutants (request-record equality; the source binding in the live digest). medium

The GPT review also confirmed the literature correction against the primary manuscripts: the LPS 2019 sample definitions (eprints.lse.ac.uk/87481, §3 pp. 9-10) and the ABJK 2022 ones (September 2021 manuscript, §3 p. 10 and §4.2.2 p. 18).

Repair (Claude Opus builder, test-first)

Commit What
7e07e7ae Failing-first command tests for H1 and M1.
f693df24 seal.json, written once at seal time before the ledger: ledger sha256, binding, completion stamps, every page sha256, attempt and page counts. core.store.recover_seal reads the ledger under the recorded digest. Each seal binds the study tree, protocol sha256 and runtime-lock sha256. The fetch start is stored and read back, never matched. Every refusal is a SealError.
5a1d4b0e Protocol text, one review_record entry for both reviews, the #358 citation, the test record.
1e3fd5a1 Follow-up after an independent Opus verification passed the repair with five low findings (four repaired; the fifth was an untracked cache directory).

A GPT builder was not used: the native Codex login reached its usage limit at 2026-10-01T00:04:04Z, before any edit.

Checks at 406ad3c5

Command Exit Result
The study's own test command (uv run … python -B -m unittest discover -s blueprints/us-equities/mover-v3/study/tests -t blueprints/us-equities/mover-v3/study) 0 Ran 395 tests, OK (376 before the repair)
python3 scripts/validate.py 0 passed
python3 scripts/evidence_manifest.py --check 0 8655 files
  • Study tree: 6cb88f03b45fad24f0fb95568b21e2834b9cdc17, the tree the protocol's run_discipline names.
  • Mutants, each killed by a named test (in review_record): study tree removed from the dry-run identity; fetch_date matched again; the plan-change check removed; recovery hashing the current ledger; request-equality checks removed; source binding omitted; page_sha256s emptied in seal_record_of.

Residuals (recorded in review_record, not repaired)

Scout and others added 7 commits October 1, 2026 00:49
…napshot path, tree content check, stamp tests

Repairs the four findings of the cross-family re-check of the round-18 repair
(GPT-6.1 Sol, effort max, Codex 0.159.2, 2026-10-01; changes-needed).

R2-1 (high). core.store._recover_seal passed the seal record's ledger_sha256
to Store.read, which compares no digest when it is given None, and
seal_record_of gave the null back. A record whose ledger_sha256 was JSON null
was adopted with its ledger unverified. The record must now hold 64 lowercase
hexadecimal characters; anything else is a SealError before the ledger is read.

R2-2 (high). core.guards.running_tree returned HEAD's study tree whenever git
status reported no change. git status reports none for a tracked file whose
index entry is marked assume-unchanged or skip-worktree (git-update-index(1);
git-ls-files(1) -t and -v), nor for one that a clean filter maps back to the
committed bytes, so the tree in every seal identity (H1) could differ from the
files that ran. running_tree now refuses an index entry with either flag, and
requires every file of the tree it returns to be a regular file whose own bytes
hash to its blob (one git hash-object --no-filters --stdin-paths call). A
symbolic-link or submodule entry is refused. The signature is unchanged.

R2-3 (medium). A regular file or a dangling symbolic link at the snapshot path
counted as unused: the caller fetched again and Store.write then failed. It is
now a SealError before any fetch.

R2-4 (medium). Two stamp controls had no test: the attempt ids in the seal
record's stamps, and Store.read's refusal of a second completion stamp for one
key and attempt. Both have one now; no code changes for this finding.

Not changed: recover_seal still returns None for a path under a regular file
and for a directory that holds only a stray ledger.jsonl.tmp or seal.json.tmp,
so the caller fetches before Store.write fails with an OSError.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ntence, re-pinned draft tree and test record

protocol-core-draft.json, after the study repair 28fafe8:

- review_record: one new round-18 entry (2026-10-01) for the cross-family
  re-check of the first repair (GPT-6.1 Sol, effort max, Codex 0.159.2,
  changes-needed) and its four findings R2-1 to R2-4. It states the repair,
  the failing-first output (Ran 46 tests, FAILED (failures=64, errors=9)),
  the direct reproductions at 406ad3c, the six mutants with the tests that
  kill them, and what is not changed. It corrects three sentences of the
  earlier round-18 entry, which is left as written.
- run_discipline.study_code.rule: one sentence on how core.guards.running_tree
  establishes the tree a run names, and that it checks once per run.
- run_discipline.study_code.draft_tree_at_this_revision, test_record and
  freeze_acceptance entry 7: draft tree b1cd175,
  407 tests (Ran 407 tests in 296.620s, OK).

Status stays draft_pending_independent_pre_outcome_review; nothing read an
outcome. manifests/evidence.json is not touched: its registration of this
file needs the coordinator's re-registration.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… without replace refs; three content-check tests

Found by a direct probe of the repair's own content check (28fafe8), before
any re-review. With a replace ref on HEAD's study tree object (git-replace(1)),
git reads another tree in its place: git status reports nothing for files that
match that other tree, no index flag is set, and git ls-tree lists its blobs
under the id of HEAD's tree. running_tree compared the files with those blobs
and returned HEAD's tree. It now reads the tree id and lists the tree with
git --no-replace-objects, so the blobs compared are the tree's own.

Tests (tests/test_guards.py): the replace-ref state, which failed first with
'AssertionError: Refused not raised' (Ran 13 tests, FAILED (failures=1)); a
tracked file name with a line break, refused before git hash-object
--stdin-paths is given any path; and git hash-object exiting with an error or
printing one hash too few, which refuses the tree. The last two cover branches
of the content check that no test reached.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r the final study tree

protocol-core-draft.json, after the study follow-up ba10683:

- review_record: the second-repair entry of 2026-10-01 (added in f1fcffe,
  not yet reviewed) now also records the replace-ref state found by a direct
  probe of the content check at 28fafe8, its fix (git --no-replace-objects),
  the three follow-up tests, the mutant runs repeated on the final tree
  (content check removed: Ran 47 tests, FAILED (failures=12); index-flag
  refusal removed: Ran 47 tests, FAILED (failures=11)) and three more mutants,
  and two more residuals: what running_tree does not check, and an untracked
  file that git status does not report when core.worktree points elsewhere.
- run_discipline.study_code.rule: the R2-2 sentence names the read without
  replace refs.
- draft_tree_at_this_revision, test_record and freeze_acceptance entry 7:
  draft tree 61f939b, 410 tests
  (Ran 410 tests in 302.526s, OK).

Relative to 406ad3c the protocol gains one review_record entry. Status stays
draft_pending_independent_pre_outcome_review; nothing read an outcome.
manifests/evidence.json is not touched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rule (F1), replace refs ignored by every guard (F2)

Two bounded items from the coordinator's ruling on the second repair.

F1 closes R2-3's class. core.store.recover_seal treats a snapshot path as
unused only if it is absent and its nearest existing ancestor is a directory,
or it is an empty directory. A path under a regular file or a dangling link,
and a directory that holds anything without a ledger (a stray
ledger.jsonl.tmp or seal.json.tmp, any other file), are refused as a
SealError before any fetch. The 'partially written' and 'exists and is not a
directory' refusals keep their wording. At 7acb694 the first two shapes made
the dry run fetch (16 calls) before Store.write failed with NotADirectoryError
or FileExistsError.

F2 moves --no-replace-objects from two calls in running_tree into
core.guards.git_command, which now builds every git command the module runs
(git, git_ok, committed_bytes, verified_merge and the running-tree checks),
so the tree diff behind transport deviations, committed_bytes and every
rev-parse, log and show read the repository's own objects (git(1);
git-replace(1)).

Tests, written first (Ran 57 tests, FAILED (failures=7, errors=8) at
7acb694): for the dry run and the live samples, a regular file or dangling
link as the snapshot root, a directory holding only a stray file, and unused
paths that are still fetched into and sealed; committed_bytes and
fetch_only_diff under a replace ref. The earlier replace-ref test is kept:
with every git command ignoring replace refs, git status itself now reports
the files as changed, so that is the refusal it asserts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s the refusal before its supporting observations

tests/test_guards.py only. With the helper's --no-replace-objects removed,
test_running_tree_refuses_files_that_match_a_replacement_of_heads_tree now
fails on running_tree itself ('Refused not raised') instead of on the
supporting git status observation that preceded it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the final study tree (F1, F2)

protocol-core-draft.json, after the study commits 89f64e2 and cd36752:

- review_record: the second-repair entry of 2026-10-01 (not yet reviewed)
  now records the coordinator's two follow-up items. F1: a snapshot path is
  unused only if it is absent under a directory or an empty directory, which
  closes the residual that entry had listed as (4). F2: every git command of
  core.guards ignores replace refs, which closes the untested option on
  rev-parse and the other guards' git calls. It gives their failing-first
  output (Ran 57 tests, FAILED (failures=7, errors=8)), the mutant runs
  repeated on the final tree and the two new ones (F1's rule removed: Ran 42
  tests, FAILED (failures=5, errors=8); the helper's protection removed: Ran
  55 tests, FAILED (failures=3)), and drops the sentences those items made
  untrue.
- run_discipline.study_code.rule: the R2-2 sentences state F2's rule.
- draft_tree_at_this_revision, test_record and freeze_acceptance entry 7:
  draft tree 5ebe141, 418 tests
  (Ran 418 tests in 282.984s, OK).

Relative to 406ad3c the protocol gains one review_record entry. Status stays
draft_pending_independent_pre_outcome_review; nothing read an outcome.
manifests/evidence.json is not touched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scout and others added 3 commits October 1, 2026 04:18
…o-filters test (T1), index-flag refusal message (T2)

The independent Opus verification of ace5648 (C7) confirmed the repair and
found one test gap and one misleading message.

T1 (test only). No test protected --no-filters in running_tree: removing it
left the module tests green (reproduced: Ran 55 tests, OK). git hash-object
--stdin-paths receives absolute paths, and git applies no anchored
repository-relative attribute pattern to an absolute path (probed on git
2.43.0: a basename or **/ pattern applies, an anchored one does not), so the
clean filter of test_running_tree_refuses_a_change_that_a_clean_filter_hides
never reached the content check. The test now uses a basename pattern,
checks that git hash-object filters the absolute path running_tree passes
unless --no-filters is given, and its docstring says so. With --no-filters
removed it now fails ('Refused not raised'; Ran 56 tests, FAILED
(failures=1)).

T2 (message only). With core.ignoreStat=true git marks every file it checks
out assume-unchanged, so a fresh, byte-identical checkout is refused. The
refusal stays; its message now says that the index marks the files
assume-unchanged or skip-worktree, so the run cannot verify them, and how to
clear the bits. Unsetting core.ignoreStat alone leaves the bits (probed), so
the message names both steps. New test, which failed first on the old
wording: test_running_tree_refuses_a_checkout_made_under_core_ignore_stat_and_says_how_to_clear_it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ication of ace5648 (T1, T2)

protocol-core-draft.json, after the study commit 50fd683:

- review_record[48] (the same unpushed second-repair entry): one sentence on
  the independent Opus verification of ace5648 and its two fixes (the
  --no-filters test gap, Ran 55 tests, OK with the option removed before
  the fix, Ran 56 tests, FAILED (failures=1) after it; and the index-flag
  refusal's message under core.ignoreStat=true). The test count and the
  wording that named cd36752 as the final tree are updated to match.
- draft_tree_at_this_revision, test_record and freeze_acceptance entry 7:
  draft tree cc7ecae, 419 tests
  (Ran 419 tests in 273.947s, OK).

Relative to 406ad3c the protocol still gains one review_record entry.
Status stays draft_pending_independent_pre_outcome_review; nothing read an
outcome. manifests/evidence.json is not touched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…protocol, last commit)

protocol-core-draft.json changed in the second repair (sha256 87a955fe..., 361469 bytes).
Only its entry changes; component_matrix.py --write and new_host_grand_list.py --write
re-run (no change). scripts/validate.py exit 0; scripts/evidence_manifest.py --check passed
(8655 files). The branch is re-merged with main at its train slot.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Cross-family re-check of the repair, second repair and verification (2026-10-01)

Head is now 56bbb2cd: ten commits on 406ad3c5 (fast-forward push), the last re-registering the protocol in manifests/evidence.json. The branch is re-merged with main at its train slot. Protocol status stays draft_pending_independent_pre_outcome_review; nothing reads an outcome.

1. GPT re-check of the first repair

  • Reviewer: gpt-6.1-sol, effort max, Codex 0.159.2, read-only, packaged lane on OmniRoute; 03:42–04:18Z.
  • Verdict: changes-needed, 4 findings, each reproduced by the reviewer and confirmed by me in source.
  • Confirmed by the same read: all seven mutants of the first repair are killed, and its three stated residuals are accurate.
# Severity Finding
R2-1 high A seal record whose ledger_sha256 is JSON null turns ledger verification off (Store.read skips the comparison when the expected digest is None). The reviewer's probe adopted an edited ledger.
R2-2 high running_tree trusted git status, which hides tracked files marked assume-unchanged or skip-worktree. A changed transport could run under the committed tree's identity and adopt an earlier seal.
R2-3 medium A regular file at a snapshot path was treated as unused: the run fetched again, then failed with NotADirectoryError, not a SealError.
R2-4 medium Two stamp controls survived as mutants: the attempt id in the seal record's stamps, and Store.read's second-stamp refusal.

2. Second repair (Claude Opus builder, test-first)

Commits.

Commits What
28fafe84, f1fcffe1 R2-1: the record's ledger_sha256 must be 64 lowercase hex before any read. R2-2: running_tree refuses index entries marked assume-unchanged or skip-worktree (git ls-files -v) and hashes every tracked study file against HEAD's blob (hash-object --no-filters). R2-3: a non-directory snapshot path is refused before any fetch. R2-4: tests for both stamp controls.
ba106834, 7acb6944 The builder found that a replace ref on HEAD's study tree fooled the content check; tree blobs are now read with --no-replace-objects.
89f64e24, cd36752e, ace56485 Two gaps of the same class, closed on my follow-up. F1: a snapshot path is unused only if it is absent under a directory or is an empty directory; anything else is refused as a SealError before a fetch. F2: one builder, git_command, runs every git command of core/guards.py with --no-replace-objects (including committed_bytes and the transport-deviation diff).
50fd6836, 850a3d9c After the verification below. T1: the --no-filters test now uses a basename attribute pattern, so git really filters the absolute path running_tree passes. T2: the index-flag refusal says how to clear the bits.
56bbb2cd Manifest re-registration.
  • Final study tree: cc7ecaecfc0207b61bd20fa4178bfa6769bbc857, named by draft_tree_at_this_revision, test_record and freeze_acceptance entry 7.
  • Protocol changes: one new review_record entry, sentences appended to run_discipline.study_code.rule, and those pins.

3. Independent verification (Claude Opus, read-only, at ace56485)

  • C1–C6 confirmed:
    • full suite 418 OK;
    • pins and scope correct;
    • each GPT finding re-probed and closed: 18 digest cases, 5 index-flag cases, 7 path cases, both stamp mutants;
    • every named mutant killed;
    • one git builder.
  • C7, new issues:
    • one medium test gap: no test protected --no-filters. The test's anchored attribute pattern never applies to the absolute paths that hash-object --stdin-paths receives. Fixed as T1, and the mutant is now killed.
    • one low: the core.ignoreStat refusal message was misleading. Fixed as T2; the refusal itself is by design.
    • one informational point.

Checks at the final tree

Command Exit Result
The study's own test command (full suite) 0 Ran 419 tests, OK (395 before this repair round)
python3 scripts/validate.py 0 passed
python3 scripts/evidence_manifest.py --check 0 8655 files

Residuals (recorded in review_record)

  • No GPT read of the second repair. The OmniRoute pool account read 92 of 100 at 08:09Z, and the remaining points went to closure-synthesis reviews that had no other second-family read. The independent Opus verification re-ran the same probes.
  • running_tree checks once per run: a file changed later in the same run is not seen, and the executable bit is not compared.
  • A core.worktree pointing at another clean copy hides an extra untracked file from git status. A changed tracked file is still refused in that state.
  • An operating-system error while Store.write writes after a fetch is not a SealError.
  • .git/info/grafts (deprecated, still read by git 2.43) can alter commit parents for the history guards. Noticed by reading, not probed.
  • Line endings depend on the repository root .gitattributes, which is outside the study tree hash. There is no false refusal in this repository's configuration.
  • Host hygiene: each full-suite run leaves mover-v3-test-gnupg-* keyrings with running gpg-agents under /var/tmp, because the atexit cleanup in tests/fixture_repo.py does not run in some processes. More than 120 remain on the trading host. This is a test-harness issue for a follow-up, not part of this PR.

seathatflowsinourveins pushed a commit that referenced this pull request Oct 1, 2026
…as done, not as "not running"

`codex_job.py wait` read `done`, then the job lock. A job that finished between the two reads was reported as
"done exit=none (the job is not running)". The runner writes `done` before it releases the lock (flock(2): the lock
goes when its last descriptor closes), so wait now reads `done` again once it sees the lock free.

- tools/sota-convergence/landscape-sweep/codex_job.py: the re-read in `wait`.
- tests/test_landscape_sweep_harness.py: `WaitFinishRaceTests` (synthetic fixture: the lock check itself completes
  the job; fails before the fix with 'done exit=none (the job is not running)' != 'done exit=0'; a job that never
  ran is still reported as not running); `RunnerCase.diagnose`, used by the web-search-modes test so a failed start
  or wait prints the command's output and the runner's files.
- docs/harness-defaults.md: two anti-pattern rows (this race; the evidence-manifest rows a hot-file re-run dropped).

Observed once: hosted run 36836458360 on PR #360 failed test_web_search_modes_are_staged_bound_and_recorded at its
wait step, with no output retained. The fixture shows the mechanism; a single-core stress loop did not reproduce
the hosted failure before the fix, so the match with that run is an inference.

Tests: python3 -m unittest tests.test_landscape_sweep_harness tests.test_adoption_docs_consistency -> 209 tests OK
(4 skipped), exit 0.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins merged commit 6706517 into main Oct 1, 2026
29 of 30 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the claude/mover-v3-round18-20260926 branch October 1, 2026 10:26
seathatflowsinourveins added a commit that referenced this pull request Oct 1, 2026
…as done, not as "not running" (#576)

`codex_job.py wait` read `done`, then the job lock. A job that finished between the two reads was reported as
"done exit=none (the job is not running)". The runner writes `done` before it releases the lock (flock(2): the lock
goes when its last descriptor closes), so wait now reads `done` again once it sees the lock free.

- tools/sota-convergence/landscape-sweep/codex_job.py: the re-read in `wait`.
- tests/test_landscape_sweep_harness.py: `WaitFinishRaceTests` (synthetic fixture: the lock check itself completes
  the job; fails before the fix with 'done exit=none (the job is not running)' != 'done exit=0'; a job that never
  ran is still reported as not running); `RunnerCase.diagnose`, used by the web-search-modes test so a failed start
  or wait prints the command's output and the runner's files.
- docs/harness-defaults.md: two anti-pattern rows (this race; the evidence-manifest rows a hot-file re-run dropped).

Observed once: hosted run 36836458360 on PR #360 failed test_web_search_modes_are_staged_bound_and_recorded at its
wait step, with no output retained. The fixture shows the mechanism; a single-core stress loop did not reproduce
the hosted failure before the fix, so the match with that run is an inference.

Tests: python3 -m unittest tests.test_landscape_sweep_harness tests.test_adoption_docs_consistency -> 209 tests OK
(4 skipped), exit 0.

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Oct 1, 2026
…tion, verbatim-checked votes (verdict-wave input) (#358)

* Trading research convergence 2026-09-26: 13 layers, GPT-6 live refutation, verbatim-checked votes

catalogs/us-equities/convergence-20260926.{json,md} covers 13 trading layers:
the 12 catalog layers plus strategy-research. They were researched from the
starred repositories, 8 awesome lists and per-layer candidate research.

- Each Claude proposal was refuted by GPT-6 (gpt-6-astra, effort max,
  read-only). Every layer has a live-search vote; 8 also have a cached one.
- One Opus mapper per layer attributed votes to candidates and winners with
  verbatim quotes. The builder kept only votes whose quote matches the
  retained vote text: 716 kept, 0 dropped. A planted-quote check confirmed
  that non-matching quotes are dropped.
- Candidates use manifest-20260926's trading candidate shape. Dispositions:
  95 survive, 40 refuted, 47 unverified.
- Failed attempts (the 401 sign-out and the stopped cached runs) are listed.

evidence/artifacts/trading-convergence-20260926/ holds the GPT-6 vote texts,
the per-layer packets and the live paper sweep, sanitized:
- paths made repository-relative, or replaced by <session-scratch> and
  <private-memory>;
- UUIDs replaced by <uuid>;
- #task- URL fragments dropped;
- 'api:' colons dropped.

The record makes no selection, pin or installation change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Point the trading convergence record at the repaired manifest-20260926 (sha a72177d162f8)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Trading convergence record: repair the pre-merge check findings (exact-span quotes, survival rule wording, round-18 status, privacy)

GPT-6 read-only pre-merge check of 260963f (needs_changes), addressed:
- Quote integrity: 223 of 716 stored quotes had typographic punctuation
  normalized. Quotes are now the source's own span (located tolerating only
  whitespace and typographic punctuation), sanitized exactly as the exported
  texts, and kept only if they occur in the published text. 716 of 716 are
  exact spans; none dropped.
- Survival rule: only position 'refutes' changes survival; 'corrects' fixes
  a fact or claim and leaves the label standing. The wording is now
  explicit.
- Round 18: the ABJK/LPS correction is described as proposed in #360, under
  review, not as done.
- Privacy: email addresses are replaced by <email>. Two claim checks that
  quoted private host memory keep their name and verdict with their text
  withheld.

The record's content and dispositions are otherwise unchanged
(95/40/47, 716 votes). validate.py, validate_catalogs.py and gitleaks are
clean; validate.py's private-content patterns find nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Trading convergence record: withhold every sentence that cites private host memory

The GPT-6 re-check of 55fa3f8 found memory-derived quotations still present
in the agents-models-workers claude_summary and in two winner notes of its
packet. The redaction had covered only the claim checks.

Every sentence that cites private host memory is now withheld in all fields
of the record and the exported packets. The shared sanitizer applies the
same rule to both, sanitizing each JSON string individually so the packets
stay valid JSON. 8 sentences are withheld.

Verification from the published files alone:
- 716 of 716 quotes are exact spans;
- all JSON files are valid;
- no email address, no private-memory marker and no validate.py
  private-content hit.
Dispositions and counts are unchanged (95/40/47).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Trading convergence record: withhold three surviving private-memory clauses; fix the companion layer count

Landing review (Claude Opus, 2026-10-01, read-only at the merged tree) found:
- medium: two mapper `reasoning` fields and the security-supply-chain summary (record and
  packet) still cited private host memory, contrary to method.mapping and the artifact
  README. Each clause is replaced by `[A clause citing private host memory is withheld.]`.
- low: companion_record.relation said the landscape manifest covers 11 of these layers with
  security-supply-chain as a foundation layer. The pinned manifest-20260926.json (sha256
  a72177d1...) lists it among its 12 trading layers, so it covers 12 of 13.

Exact-byte edits, because the builder is not retained; no quote and no vote file changed.
An independent quote check finds 716 of 716 quotes in the retained texts (578 in vote files,
138 in packet claim checks) before and after, and reports a planted altered quote and a
planted fabricated quote as missing. method.mapping and the artifact README now state the
withholding rule and the hand edit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Trading convergence record: mark the NautilusTrader winner as cross-family disputed; neutral name for one redacted check

Cross-family landing review (gpt-6.1-sol, effort max, Codex 0.159.2, read-only, packaged
lane on OmniRoute, 2026-10-01 04:21-04:50Z) found two medium defects:
- Both GPT-6 votes on backtesting-engine's nautilustrader winner call its earns_it rationale
  circular and say comparative superiority is unproven, yet were mapped as corrects. Votes
  of the same substance on foundation-ai-memory are mapped refutes. The two votes are now
  refutes (each reasoning carries a dated note), the winner is cross_family_disputed, and
  the .md table row names it. No candidate disposition or count changes.
- One redacted claim check kept a name that cited private host memory. The name is now
  neutral, and the artifact README says names do not cite it.

The review confirmed: 716 of 716 quotes in the right layer, lens and source container (with
planted negative controls), all reconstructible counts, every candidate's survival rule,
the 13 table rows, the manifest hashes for exactly the 39 files, and every relative link.

Exact-byte edits (the builder is not retained); no quote or vote file changed. The four
changed files are re-registered in manifests/evidence.json. Checks on this tree:
validate.py, validate_catalogs.py and evidence_manifest.py --check exit 0; tests.test_catalogs
and tests.test_evidence_manifest Ran 52, OK; independent quote check 716 of 716.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:trading Trading lane: sim, paper and live north star

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant