Repository navigation
Mover v3 round 18: pre-freeze residual coverage, seal recovery for dry runs and live samples, ABJK/LPS literature correction — not frozen - #360
Conversation
…y runs and live samples, literature correction Residuals the round-17 verifiers and the GPT-6 re-review named: - R1-R4, coverage only. Eight tests pass on unchanged 803bc35 production code: - the H3-c split on the last overnight leg, and a split and a spin-off on different legs; - a name_change dated exactly on asof t, for every issuer-record consumer and the terminal-zero rename diagnostic; - the exact terminal_actions refusal reason; - a direct test of the no-valid-page-calibration guard. - R5: count_only.dry_run and transport_check.seal_live_samples adopt the recover-if-matching pattern of core.holdout.collect. An existing snapshot is validated with Store.read; its binding and complete request/stamp set must match the new request; its sha256 is reused without fetching or writing. Partial, mismatched, unbound or altered snapshots refuse. Eight tests are in tests/test_seal_recovery.py. Reverting only the two production files gives 30 failures and 1 error. The dry-run local variable is named seal_identity so it does not shadow the imported core.identity module. - Literature: item_set.items[4].why_kept described the LPS 2019 and ABJK 2022 samples together as 'large, liquid names'. Re-verified from the manuscripts: LPS excludes microcaps (below $5 or the bottom NYSE size quintile) and opens at the first-half-hour VWAP; ABJK keeps common stocks above $1 (excluding financials and utilities), with its effect in all but the largest decile. The test stays two-sided; no rule changes. Two round-18 review_record entries. Study test_command: 376 tests OK. The protocol stays draft_pending_independent_pre_outcome_review, with frozen_before_outcomes false. R1-R5 were implemented by GPT-6 (gpt-6-astra, effort max) test-first and verified by the coordinator; the literature correction was found by the GPT-6 paper sweep and verified by the coordinator. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
GPT-6 cross-family pre-outcome review of af38dce (
The independent Claude review is still running. |
|
The independent Claude review of round 18 was interrupted by a Claude session limit (HTTP 429, resets 22:10Z). R1–R5 were written by GPT-6, so the cross-family review must come from Claude; no other reviewer substitutes. This PR waits for that review before merging. Nothing is time-critical: the freeze is blocked on the market-data dry run in any case. |
…t-span quotes, survival rule wording, round-18 status, privacy) GPT-6 read-only pre-merge check of 260963f (needs_changes), addressed: - Quote integrity: 223 of 716 stored quotes had typographic punctuation normalized. Quotes are now the source's own span (located tolerating only whitespace and typographic punctuation), sanitized exactly as the exported texts, and kept only if they occur in the published text. 716 of 716 are exact spans; none dropped. - Survival rule: only position 'refutes' changes survival; 'corrects' fixes a fact or claim and leaves the label standing. The wording is now explicit. - Round 18: the ABJK/LPS correction is described as proposed in #360, under review, not as done. - Privacy: email addresses are replaced by <email>. Two claim checks that quoted private host memory keep their name and verdict with their text withheld. The record's content and dispositions are otherwise unchanged (95/40/47, 716 votes). validate.py, validate_catalogs.py and gitleaks are clean; validate.py's private-content patterns find nothing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Two command-level regressions for the dry-run seal recovery, through run.py dry-run against the fixture repository. Each run's output step fails after the seal, so run.py's finally block writes a failed end line and the next run opens a new start line. - H1: tree B, which changes only study/fetch/ and so keeps the same request plan, must not adopt tree A's seal. At af38dce it does: "AssertionError: SealError not raised" at test_seal_recovery.py:239. - M1: the same tree retrying on the next UTC day must adopt its seal and read back the first run's start. At af38dce the fetch date is matched: "core.store.SealError: dry-run snapshot binding differs from the requested dry run" (count_only.py:486). The binding case that asserted a fetch-date refusal (the wrong rule, the review's M1) is dropped. Module run at af38dce: Ran 10 tests, FAILED (failures=1, errors=1). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eal recovery
Resolves the round-18 pre-outcome reviews of the two recoverable seals
(count_only.dry_run, transport_check.seal_live_samples), following
core.holdout.collect's identity shape.
- G-H: Store.write(seal_record=True) publishes seal.json once, before the
ledger. It holds the ledger sha256, the binding, every completion stamp
[key, attempt, status, pages], every page sha256, and the attempt and
page counts. core.store.recover_seal reads the ledger under the recorded
digest, never a digest of the ledger being checked, and the snapshot
must give the record back exactly. Store.read checks each stamp's page
count against its page events.
- H1: each seal binds the running study tree, protocol sha256 and
runtime-lock sha256 (runner.seal_run_identity, validated by
store.checked_run_identity). The dry-run identity is {study_tree,
protocol_sha256, runtime_lock_sha256, sessions, symbols, requests}; live
samples add the same three fields to label, seed, source and requests.
- M1: the dry run's fetch start (utc_start) is bound beside the identity,
read back into the output as fetch_utc_start, and never matched.
- G-M: separate checks for the identity, the request plan, and the ledger's
request records against the requests the binding names.
- L1: attempt stamps are compared exactly through the seal record.
- L2: every recovery refusal is a SealError (zlib.error, TypeError,
AttributeError and OSError included), naming the dry run or the label.
- L3: run-wide changes are sealed into a mixed live root, so each label's
own refusal is asserted.
Tests: tests/test_seal_recovery.py, 15 new unit cases. The commit-1
command tests now pass. Callers in test_count_only.py and test_fetch.py
pass the run identity. Study suite: Ran 393 tests, OK.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…test record
- freeze_preconditions (native dry run) and run_discipline.transport_deviations
now describe recovery through core.store.recover_seal:
- verification against the seal record;
- a binding to the running study tree, protocol sha256 and runtime-lock
sha256 plus the caller's inputs and request plan;
- the dry run's fetch start read back as fetch_utc_start, never matched;
- a new snapshot root for any other tree, protocol or runtime lock.
- review_record: one round-18 entry dated 2026-10-01 for the Claude
(H1, M1, L1-L5) and GPT (G-H, G-M) reviews. It records the failing-first
output, the reproduction at af38dce, the mutants with their killing
tests, the record's limit (a write-once file, not an external anchor),
what is not changed (holdout.collect; orphan adoption), and the GPT
review's confirmation of the LPS 2019 and ABJK 2022 sample definitions.
- L4: the round-18 literature entry cites PR #358 (not yet on main),
7e29448:evidence/artifacts/trading-convergence-20260926/papers-strategy.md.
- L5: draft_tree_at_this_revision, test_record and freeze_acceptance[7]
name tree 20a3dab (fetch tree
unchanged) with 393 tests. test_record.command is the command actually
run; command_note explains the environment prefix.
The study tree is unchanged by this commit.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…qualified recovery rules, review record An independent verification of the round-18 repair passed it with five low findings; four are repaired. The fourth, a stray git-ignored scripts/__pycache__ in the review worktree, was deleted and changes no tracked file. - Protocol: the dry-run sentence of freeze_preconditions and run_discipline.transport_deviations no longer say that an altered snapshot refuses outright. A snapshot without its seal record is refused, and a partial, unbound, mismatched or altered one is refused when its seal.json is intact. seal.json is a write-once record inside the snapshot directory, not an external anchor: a consistent rewrite of both the ledger and seal.json that keeps the binding, the request plan and the requests is adopted. - tests/test_seal_recovery.py: test_a_resealed_page_replacement_refuses (dry run, and each live-sample label) replaces a page with its ledger digests and reseals only ledger_sha256 in seal.json. It asserts 'the snapshot differs from its seal record', no transport call and unchanged bytes. The probe mutant that empties page_sha256s in core.store.seal_record_of passed all 393 tests at 5a1d4b0; it now fails this test's three cases and nothing else (Ran 395 tests, FAILED (failures=3)). - review_record, round-18 repair entry: one sentence names the three diff items it omitted (test_record.command with its command_note, Store.read's second-stamp refusal for every snapshot type, Store.write's refusal of a directory holding seal.json). Three residuals are recorded: the seal.json limit; core.holdout.collect reading its seal under a digest of its own ledger (predates this round); orphan-output adoption verifying seals against the sha256 values the output body names, not seal.json. - run_discipline.study_code.draft_tree_at_this_revision, test_record and the freeze_acceptance evidence name the new study tree 6cb88f0 and 395 tests (fetch tree f288cf2 unchanged). Study suite: Ran 395 tests, OK. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
manifests/evidence.json is main's copy (3361b34) with protocol-core-draft.json re-registered through scripts/host_receipts.register_file; component_matrix.py --write and new_host_grand_list.py --write re-run (no change). scripts/validate.py and scripts/evidence_manifest.py --check exit 0 on this tree. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Independent pre-outcome reviews and repair (trading lane, 2026-10-01)Head is now Reviews of
|
| Review | Finding | Severity |
|---|---|---|
| Claude Opus, effort max, read-only (2026-09-30) | H1. The dry-run seal binding had no study tree, so a later tree could adopt an earlier tree's sealed snapshot without fetching. | high |
M1. fetch_date was part of the matched identity but came from the current run, so a rerun on a later UTC day was refused for good. |
medium | |
L1–L5. Attempt-stamp comparison; refusals that escaped SealError; a subTest that never reached its label; a citation of a file that is only on #358's branch; a stale tree and test count in run_discipline. |
low | |
GPT gpt-6.1-sol, effort ultra, native Codex 0.159.2, read-only (2026-10-01 00:03Z) |
G-H. Recovery trusted a digest it computed from the files it was checking. In the reviewer's synthetic run, a replaced page with matching ledger digests was adopted and turned passes=false into passes=true. |
high |
| G-M. Two recovery controls survived the tests as mutants (request-record equality; the source binding in the live digest). | medium |
The GPT review also confirmed the literature correction against the primary manuscripts: the LPS 2019 sample definitions (eprints.lse.ac.uk/87481, §3 pp. 9-10) and the ABJK 2022 ones (September 2021 manuscript, §3 p. 10 and §4.2.2 p. 18).
Repair (Claude Opus builder, test-first)
| Commit | What |
|---|---|
7e07e7ae |
Failing-first command tests for H1 and M1. |
f693df24 |
seal.json, written once at seal time before the ledger: ledger sha256, binding, completion stamps, every page sha256, attempt and page counts. core.store.recover_seal reads the ledger under the recorded digest. Each seal binds the study tree, protocol sha256 and runtime-lock sha256. The fetch start is stored and read back, never matched. Every refusal is a SealError. |
5a1d4b0e |
Protocol text, one review_record entry for both reviews, the #358 citation, the test record. |
1e3fd5a1 |
Follow-up after an independent Opus verification passed the repair with five low findings (four repaired; the fifth was an untracked cache directory). |
A GPT builder was not used: the native Codex login reached its usage limit at 2026-10-01T00:04:04Z, before any edit.
Checks at 406ad3c5
| Command | Exit | Result |
|---|---|---|
The study's own test command (uv run … python -B -m unittest discover -s blueprints/us-equities/mover-v3/study/tests -t blueprints/us-equities/mover-v3/study) |
0 | Ran 395 tests, OK (376 before the repair) |
python3 scripts/validate.py |
0 | passed |
python3 scripts/evidence_manifest.py --check |
0 | 8655 files |
- Study tree:
6cb88f03b45fad24f0fb95568b21e2834b9cdc17, the tree the protocol'srun_disciplinenames. - Mutants, each killed by a named test (in
review_record): study tree removed from the dry-run identity;fetch_datematched again; the plan-change check removed; recovery hashing the current ledger; request-equality checks removed; source binding omitted;page_sha256semptied inseal_record_of.
Residuals (recorded in review_record, not repaired)
seal.jsonis a write-once file inside the snapshot directory, not an external anchor. A consistent rewrite of both the ledger andseal.jsonthat keeps the binding, the request plan and the requests is adopted.core.holdout.collectreads its seal under a digest of its own ledger. This predates round 18.- Orphan-output adoption verifies seals against the sha256 values the output body names, not against
seal.json. - No GPT re-check of the repair. The shared pool is exhausted and native Codex runs are paused for Gate A. The repair's second read is the Opus verification above.
- The round-18 literature entry cites
papers-strategy.mdat Trading research convergence 2026-09-26: 13 layers, GPT-6 live refutation, verbatim-checked votes (verdict-wave input) #358's commit7e29448d. Trading research convergence 2026-09-26: 13 layers, GPT-6 live refutation, verbatim-checked votes (verdict-wave input) #358 lands before this PR in the merge train; the citation is updated to the repository path at this PR's slot.
…napshot path, tree content check, stamp tests Repairs the four findings of the cross-family re-check of the round-18 repair (GPT-6.1 Sol, effort max, Codex 0.159.2, 2026-10-01; changes-needed). R2-1 (high). core.store._recover_seal passed the seal record's ledger_sha256 to Store.read, which compares no digest when it is given None, and seal_record_of gave the null back. A record whose ledger_sha256 was JSON null was adopted with its ledger unverified. The record must now hold 64 lowercase hexadecimal characters; anything else is a SealError before the ledger is read. R2-2 (high). core.guards.running_tree returned HEAD's study tree whenever git status reported no change. git status reports none for a tracked file whose index entry is marked assume-unchanged or skip-worktree (git-update-index(1); git-ls-files(1) -t and -v), nor for one that a clean filter maps back to the committed bytes, so the tree in every seal identity (H1) could differ from the files that ran. running_tree now refuses an index entry with either flag, and requires every file of the tree it returns to be a regular file whose own bytes hash to its blob (one git hash-object --no-filters --stdin-paths call). A symbolic-link or submodule entry is refused. The signature is unchanged. R2-3 (medium). A regular file or a dangling symbolic link at the snapshot path counted as unused: the caller fetched again and Store.write then failed. It is now a SealError before any fetch. R2-4 (medium). Two stamp controls had no test: the attempt ids in the seal record's stamps, and Store.read's refusal of a second completion stamp for one key and attempt. Both have one now; no code changes for this finding. Not changed: recover_seal still returns None for a path under a regular file and for a directory that holds only a stray ledger.jsonl.tmp or seal.json.tmp, so the caller fetches before Store.write fails with an OSError. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ntence, re-pinned draft tree and test record protocol-core-draft.json, after the study repair 28fafe8: - review_record: one new round-18 entry (2026-10-01) for the cross-family re-check of the first repair (GPT-6.1 Sol, effort max, Codex 0.159.2, changes-needed) and its four findings R2-1 to R2-4. It states the repair, the failing-first output (Ran 46 tests, FAILED (failures=64, errors=9)), the direct reproductions at 406ad3c, the six mutants with the tests that kill them, and what is not changed. It corrects three sentences of the earlier round-18 entry, which is left as written. - run_discipline.study_code.rule: one sentence on how core.guards.running_tree establishes the tree a run names, and that it checks once per run. - run_discipline.study_code.draft_tree_at_this_revision, test_record and freeze_acceptance entry 7: draft tree b1cd175, 407 tests (Ran 407 tests in 296.620s, OK). Status stays draft_pending_independent_pre_outcome_review; nothing read an outcome. manifests/evidence.json is not touched: its registration of this file needs the coordinator's re-registration. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… without replace refs; three content-check tests Found by a direct probe of the repair's own content check (28fafe8), before any re-review. With a replace ref on HEAD's study tree object (git-replace(1)), git reads another tree in its place: git status reports nothing for files that match that other tree, no index flag is set, and git ls-tree lists its blobs under the id of HEAD's tree. running_tree compared the files with those blobs and returned HEAD's tree. It now reads the tree id and lists the tree with git --no-replace-objects, so the blobs compared are the tree's own. Tests (tests/test_guards.py): the replace-ref state, which failed first with 'AssertionError: Refused not raised' (Ran 13 tests, FAILED (failures=1)); a tracked file name with a line break, refused before git hash-object --stdin-paths is given any path; and git hash-object exiting with an error or printing one hash too few, which refuses the tree. The last two cover branches of the content check that no test reached. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r the final study tree protocol-core-draft.json, after the study follow-up ba10683: - review_record: the second-repair entry of 2026-10-01 (added in f1fcffe, not yet reviewed) now also records the replace-ref state found by a direct probe of the content check at 28fafe8, its fix (git --no-replace-objects), the three follow-up tests, the mutant runs repeated on the final tree (content check removed: Ran 47 tests, FAILED (failures=12); index-flag refusal removed: Ran 47 tests, FAILED (failures=11)) and three more mutants, and two more residuals: what running_tree does not check, and an untracked file that git status does not report when core.worktree points elsewhere. - run_discipline.study_code.rule: the R2-2 sentence names the read without replace refs. - draft_tree_at_this_revision, test_record and freeze_acceptance entry 7: draft tree 61f939b, 410 tests (Ran 410 tests in 302.526s, OK). Relative to 406ad3c the protocol gains one review_record entry. Status stays draft_pending_independent_pre_outcome_review; nothing read an outcome. manifests/evidence.json is not touched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rule (F1), replace refs ignored by every guard (F2) Two bounded items from the coordinator's ruling on the second repair. F1 closes R2-3's class. core.store.recover_seal treats a snapshot path as unused only if it is absent and its nearest existing ancestor is a directory, or it is an empty directory. A path under a regular file or a dangling link, and a directory that holds anything without a ledger (a stray ledger.jsonl.tmp or seal.json.tmp, any other file), are refused as a SealError before any fetch. The 'partially written' and 'exists and is not a directory' refusals keep their wording. At 7acb694 the first two shapes made the dry run fetch (16 calls) before Store.write failed with NotADirectoryError or FileExistsError. F2 moves --no-replace-objects from two calls in running_tree into core.guards.git_command, which now builds every git command the module runs (git, git_ok, committed_bytes, verified_merge and the running-tree checks), so the tree diff behind transport deviations, committed_bytes and every rev-parse, log and show read the repository's own objects (git(1); git-replace(1)). Tests, written first (Ran 57 tests, FAILED (failures=7, errors=8) at 7acb694): for the dry run and the live samples, a regular file or dangling link as the snapshot root, a directory holding only a stray file, and unused paths that are still fetched into and sealed; committed_bytes and fetch_only_diff under a replace ref. The earlier replace-ref test is kept: with every git command ignoring replace refs, git status itself now reports the files as changed, so that is the refusal it asserts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s the refusal before its supporting observations
tests/test_guards.py only. With the helper's --no-replace-objects removed,
test_running_tree_refuses_files_that_match_a_replacement_of_heads_tree now
fails on running_tree itself ('Refused not raised') instead of on the
supporting git status observation that preceded it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the final study tree (F1, F2) protocol-core-draft.json, after the study commits 89f64e2 and cd36752: - review_record: the second-repair entry of 2026-10-01 (not yet reviewed) now records the coordinator's two follow-up items. F1: a snapshot path is unused only if it is absent under a directory or an empty directory, which closes the residual that entry had listed as (4). F2: every git command of core.guards ignores replace refs, which closes the untested option on rev-parse and the other guards' git calls. It gives their failing-first output (Ran 57 tests, FAILED (failures=7, errors=8)), the mutant runs repeated on the final tree and the two new ones (F1's rule removed: Ran 42 tests, FAILED (failures=5, errors=8); the helper's protection removed: Ran 55 tests, FAILED (failures=3)), and drops the sentences those items made untrue. - run_discipline.study_code.rule: the R2-2 sentences state F2's rule. - draft_tree_at_this_revision, test_record and freeze_acceptance entry 7: draft tree 5ebe141, 418 tests (Ran 418 tests in 282.984s, OK). Relative to 406ad3c the protocol gains one review_record entry. Status stays draft_pending_independent_pre_outcome_review; nothing read an outcome. manifests/evidence.json is not touched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…o-filters test (T1), index-flag refusal message (T2) The independent Opus verification of ace5648 (C7) confirmed the repair and found one test gap and one misleading message. T1 (test only). No test protected --no-filters in running_tree: removing it left the module tests green (reproduced: Ran 55 tests, OK). git hash-object --stdin-paths receives absolute paths, and git applies no anchored repository-relative attribute pattern to an absolute path (probed on git 2.43.0: a basename or **/ pattern applies, an anchored one does not), so the clean filter of test_running_tree_refuses_a_change_that_a_clean_filter_hides never reached the content check. The test now uses a basename pattern, checks that git hash-object filters the absolute path running_tree passes unless --no-filters is given, and its docstring says so. With --no-filters removed it now fails ('Refused not raised'; Ran 56 tests, FAILED (failures=1)). T2 (message only). With core.ignoreStat=true git marks every file it checks out assume-unchanged, so a fresh, byte-identical checkout is refused. The refusal stays; its message now says that the index marks the files assume-unchanged or skip-worktree, so the run cannot verify them, and how to clear the bits. Unsetting core.ignoreStat alone leaves the bits (probed), so the message names both steps. New test, which failed first on the old wording: test_running_tree_refuses_a_checkout_made_under_core_ignore_stat_and_says_how_to_clear_it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ication of ace5648 (T1, T2) protocol-core-draft.json, after the study commit 50fd683: - review_record[48] (the same unpushed second-repair entry): one sentence on the independent Opus verification of ace5648 and its two fixes (the --no-filters test gap, Ran 55 tests, OK with the option removed before the fix, Ran 56 tests, FAILED (failures=1) after it; and the index-flag refusal's message under core.ignoreStat=true). The test count and the wording that named cd36752 as the final tree are updated to match. - draft_tree_at_this_revision, test_record and freeze_acceptance entry 7: draft tree cc7ecae, 419 tests (Ran 419 tests in 273.947s, OK). Relative to 406ad3c the protocol still gains one review_record entry. Status stays draft_pending_independent_pre_outcome_review; nothing read an outcome. manifests/evidence.json is not touched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…protocol, last commit) protocol-core-draft.json changed in the second repair (sha256 87a955fe..., 361469 bytes). Only its entry changes; component_matrix.py --write and new_host_grand_list.py --write re-run (no change). scripts/validate.py exit 0; scripts/evidence_manifest.py --check passed (8655 files). The branch is re-merged with main at its train slot. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Cross-family re-check of the repair, second repair and verification (2026-10-01)Head is now 1. GPT re-check of the first repair
2. Second repair (Claude Opus builder, test-first)Commits.
3. Independent verification (Claude Opus, read-only, at
|
| Command | Exit | Result |
|---|---|---|
| The study's own test command (full suite) | 0 | Ran 419 tests, OK (395 before this repair round) |
python3 scripts/validate.py |
0 | passed |
python3 scripts/evidence_manifest.py --check |
0 | 8655 files |
Residuals (recorded in review_record)
- No GPT read of the second repair. The OmniRoute pool account read 92 of 100 at 08:09Z, and the remaining points went to closure-synthesis reviews that had no other second-family read. The independent Opus verification re-ran the same probes.
running_treechecks once per run: a file changed later in the same run is not seen, and the executable bit is not compared.- A
core.worktreepointing at another clean copy hides an extra untracked file fromgit status. A changed tracked file is still refused in that state. - An operating-system error while
Store.writewrites after a fetch is not aSealError. .git/info/grafts(deprecated, still read by git 2.43) can alter commit parents for the history guards. Noticed by reading, not probed.- Line endings depend on the repository root
.gitattributes, which is outside the study tree hash. There is no false refusal in this repository's configuration. - Host hygiene: each full-suite run leaves
mover-v3-test-gnupg-*keyrings with running gpg-agents under/var/tmp, because theatexitcleanup intests/fixture_repo.pydoes not run in some processes. More than 120 remain on the trading host. This is a test-harness issue for a follow-up, not part of this PR.
…as done, not as "not running" `codex_job.py wait` read `done`, then the job lock. A job that finished between the two reads was reported as "done exit=none (the job is not running)". The runner writes `done` before it releases the lock (flock(2): the lock goes when its last descriptor closes), so wait now reads `done` again once it sees the lock free. - tools/sota-convergence/landscape-sweep/codex_job.py: the re-read in `wait`. - tests/test_landscape_sweep_harness.py: `WaitFinishRaceTests` (synthetic fixture: the lock check itself completes the job; fails before the fix with 'done exit=none (the job is not running)' != 'done exit=0'; a job that never ran is still reported as not running); `RunnerCase.diagnose`, used by the web-search-modes test so a failed start or wait prints the command's output and the runner's files. - docs/harness-defaults.md: two anti-pattern rows (this race; the evidence-manifest rows a hot-file re-run dropped). Observed once: hosted run 36836458360 on PR #360 failed test_web_search_modes_are_staged_bound_and_recorded at its wait step, with no output retained. The fixture shows the mechanism; a single-core stress loop did not reproduce the hosted failure before the fix, so the match with that run is an inference. Tests: python3 -m unittest tests.test_landscape_sweep_harness tests.test_adoption_docs_consistency -> 209 tests OK (4 skipped), exit 0. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…as done, not as "not running" (#576) `codex_job.py wait` read `done`, then the job lock. A job that finished between the two reads was reported as "done exit=none (the job is not running)". The runner writes `done` before it releases the lock (flock(2): the lock goes when its last descriptor closes), so wait now reads `done` again once it sees the lock free. - tools/sota-convergence/landscape-sweep/codex_job.py: the re-read in `wait`. - tests/test_landscape_sweep_harness.py: `WaitFinishRaceTests` (synthetic fixture: the lock check itself completes the job; fails before the fix with 'done exit=none (the job is not running)' != 'done exit=0'; a job that never ran is still reported as not running); `RunnerCase.diagnose`, used by the web-search-modes test so a failed start or wait prints the command's output and the runner's files. - docs/harness-defaults.md: two anti-pattern rows (this race; the evidence-manifest rows a hot-file re-run dropped). Observed once: hosted run 36836458360 on PR #360 failed test_web_search_modes_are_staged_bound_and_recorded at its wait step, with no output retained. The fixture shows the mechanism; a single-core stress loop did not reproduce the hosted failure before the fix, so the match with that run is an inference. Tests: python3 -m unittest tests.test_landscape_sweep_harness tests.test_adoption_docs_consistency -> 209 tests OK (4 skipped), exit 0. Co-authored-by: Scout <scout@local> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…tion, verbatim-checked votes (verdict-wave input) (#358) * Trading research convergence 2026-09-26: 13 layers, GPT-6 live refutation, verbatim-checked votes catalogs/us-equities/convergence-20260926.{json,md} covers 13 trading layers: the 12 catalog layers plus strategy-research. They were researched from the starred repositories, 8 awesome lists and per-layer candidate research. - Each Claude proposal was refuted by GPT-6 (gpt-6-astra, effort max, read-only). Every layer has a live-search vote; 8 also have a cached one. - One Opus mapper per layer attributed votes to candidates and winners with verbatim quotes. The builder kept only votes whose quote matches the retained vote text: 716 kept, 0 dropped. A planted-quote check confirmed that non-matching quotes are dropped. - Candidates use manifest-20260926's trading candidate shape. Dispositions: 95 survive, 40 refuted, 47 unverified. - Failed attempts (the 401 sign-out and the stopped cached runs) are listed. evidence/artifacts/trading-convergence-20260926/ holds the GPT-6 vote texts, the per-layer packets and the live paper sweep, sanitized: - paths made repository-relative, or replaced by <session-scratch> and <private-memory>; - UUIDs replaced by <uuid>; - #task- URL fragments dropped; - 'api:' colons dropped. The record makes no selection, pin or installation change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Point the trading convergence record at the repaired manifest-20260926 (sha a72177d162f8) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Trading convergence record: repair the pre-merge check findings (exact-span quotes, survival rule wording, round-18 status, privacy) GPT-6 read-only pre-merge check of 260963f (needs_changes), addressed: - Quote integrity: 223 of 716 stored quotes had typographic punctuation normalized. Quotes are now the source's own span (located tolerating only whitespace and typographic punctuation), sanitized exactly as the exported texts, and kept only if they occur in the published text. 716 of 716 are exact spans; none dropped. - Survival rule: only position 'refutes' changes survival; 'corrects' fixes a fact or claim and leaves the label standing. The wording is now explicit. - Round 18: the ABJK/LPS correction is described as proposed in #360, under review, not as done. - Privacy: email addresses are replaced by <email>. Two claim checks that quoted private host memory keep their name and verdict with their text withheld. The record's content and dispositions are otherwise unchanged (95/40/47, 716 votes). validate.py, validate_catalogs.py and gitleaks are clean; validate.py's private-content patterns find nothing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Trading convergence record: withhold every sentence that cites private host memory The GPT-6 re-check of 55fa3f8 found memory-derived quotations still present in the agents-models-workers claude_summary and in two winner notes of its packet. The redaction had covered only the claim checks. Every sentence that cites private host memory is now withheld in all fields of the record and the exported packets. The shared sanitizer applies the same rule to both, sanitizing each JSON string individually so the packets stay valid JSON. 8 sentences are withheld. Verification from the published files alone: - 716 of 716 quotes are exact spans; - all JSON files are valid; - no email address, no private-memory marker and no validate.py private-content hit. Dispositions and counts are unchanged (95/40/47). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Trading convergence record: withhold three surviving private-memory clauses; fix the companion layer count Landing review (Claude Opus, 2026-10-01, read-only at the merged tree) found: - medium: two mapper `reasoning` fields and the security-supply-chain summary (record and packet) still cited private host memory, contrary to method.mapping and the artifact README. Each clause is replaced by `[A clause citing private host memory is withheld.]`. - low: companion_record.relation said the landscape manifest covers 11 of these layers with security-supply-chain as a foundation layer. The pinned manifest-20260926.json (sha256 a72177d1...) lists it among its 12 trading layers, so it covers 12 of 13. Exact-byte edits, because the builder is not retained; no quote and no vote file changed. An independent quote check finds 716 of 716 quotes in the retained texts (578 in vote files, 138 in packet claim checks) before and after, and reports a planted altered quote and a planted fabricated quote as missing. method.mapping and the artifact README now state the withholding rule and the hand edit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Trading convergence record: mark the NautilusTrader winner as cross-family disputed; neutral name for one redacted check Cross-family landing review (gpt-6.1-sol, effort max, Codex 0.159.2, read-only, packaged lane on OmniRoute, 2026-10-01 04:21-04:50Z) found two medium defects: - Both GPT-6 votes on backtesting-engine's nautilustrader winner call its earns_it rationale circular and say comparative superiority is unproven, yet were mapped as corrects. Votes of the same substance on foundation-ai-memory are mapped refutes. The two votes are now refutes (each reasoning carries a dated note), the winner is cross_family_disputed, and the .md table row names it. No candidate disposition or count changes. - One redacted claim check kept a name that cited private host memory. The name is now neutral, and the artifact README says names do not cite it. The review confirmed: 716 of 716 quotes in the right layer, lens and source container (with planted negative controls), all reconstructible counts, every candidate's survival rule, the 13 table rows, the manifest hashes for exactly the 39 files, and every relative link. Exact-byte edits (the builder is not retained); no quote or vote file changed. The four changed files are re-registered in manifests/evidence.json. Checks on this tree: validate.py, validate_catalogs.py and evidence_manifest.py --check exit 0; tests.test_catalogs and tests.test_evidence_manifest Ran 52, OK; independent quote check 716 of 716. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Status (2026-09-26): pre-freeze residuals, not frozen; independent reviews in progress
Round 17 (#351, GPT-6
ready_to_freeze) left residuals: coverage the verifiers named, and one fail-closed usability gap. This PR closes them and makes one literature correction. The protocol staysdraft_pending_independent_pre_outcome_reviewwithfrozen_before_outcomes: false. No outcome data, sealed file or market data was opened.Round 18
name_changeexactly at asof t, for everyCtx.issuer_recordsconsumer and the terminal-zero rename diagnosticterminal_actionsrefusal reason and the planner boundarycount_only.dry_runandtransport_check.seal_live_samplesrecover a matching sealed snapshot and refuse partial, mismatched, unbound or altered ones, with no refetch or overwrite (theholdout.collectpattern)test_seal_recovery.py; reverting the two files gives 30 failures and 1 erroritem_set.items[4].why_keptdescribed the LPS 2019 and ABJK 2022 samples together as "large, liquid names"evidence/artifacts/trading-convergence-20260926/papers-strategy.md, landing in Trading research convergence 2026-09-26: 13 layers, GPT-6 live refutation, verbatim-checked votes (verdict-wave input) #358).The two new round-18
review_recordentries cover R1–R5 and the literature correction.Checks
scripts/validate.pypassed. The only evidence change is the protocol pin; study code and tests are not hash-listed.gpt-6-astra, effort max) implemented R1–R5 test-first. The coordinator verified them and renamed one local variable indry_runthat shadowed the importedcore.identitymodule. The coordinator made the literature correction.Path to a freeze (unchanged)
freeze_preconditionshold.SOTA sources
core.holdout.collectrecover-if-matching at 803bc35 (this repository), the pattern R5 follows🤖 Generated with Claude Code