Skip to content

Mover studies: preregistered v1 and v2 (no pass in any family), data corrections D1-D6 - #162

Merged
seathatflowsinourveins merged 10 commits into
mainfrom
claude/mover-early-entry-20260924
Sep 24, 2026
Merged

seathatflowsinourveins merged 10 commits into
mainfrom
claude/mover-early-entry-20260924

Conversation

@seathatflowsinourveins

@seathatflowsinourveins seathatflowsinourveins commented Sep 24, 2026 •

Copy link
Copy Markdown
Owner

Summary

Preregistered, reproducible tests of whether long-only rules can pre-position in or enter US-equity extreme movers early, net of quote-measured costs, with a preregistered auto-leverage schedule (rungs 1/2/4 x regime x drawdown factors, cooldown, capacity caps).

  • v1 (protocol.json, frozen 4e57667a): 8 clock times x thresholds x volume floors x news x 4 exits = 768 rule-exits. No development pass; all 704 rule-exits with >= 200 trades are negative net (median of rule-exit means -2.9%); validation similar (-2.7%). Holdout not read. Results reproduced byte for byte (sha256 9d8c3d5b...).
  • v2 (protocol-v2.json, frozen a4a2f682 before any v1 or v2 outcome): E first-cross (earliest-stage) entries (216), F opening-range / pre-market-high / VWAP-reclaim setups (16), P pre-positioning at the prior close and day-2 continuation (12). No development pass in any family: E all negative even before costs; F best -0.34% net; P best -1.0% net, though its volume-breakout score holds 11.7% of next-day >= +100% movers (base rate 0.004%). Results sha256 3334962f..., byte-identical reruns.
  • Data corrections before outcomes (deviations.json): D2 bars and auctions re-collected with the candidate list's naming date (asof = session lost 16.8% of candidates); D3 Alpaca daily bars are regular-session bars, so a pre-market scan (36,185 pages, all 200) added 40,354 pre-market mover days; D4 SPY window; D1 an ordering slip (deleted unread); D5 the independent v1 pre-outcome review (10 serious findings, resolved). D6 (post-outcome, labelled): degree-tier capture and replication also without basis-uncertain rows; trade metrics and selection unchanged.
  • Two independent pre-outcome reviews (v1: 3 lenses + refuters; v2 plus a v1 result audit: 3 lenses + refuters), all serious findings resolved before the outcome they could affect (clarifications C1-C32, V1-V20).
  • Live scanner (mover_scan.py, full market in ~6 s, 33 requests) reuses the study's rules; paper_compare.py scores paper fills against the historical fill model.

Evidence

  • evidence/cost-table-run-v1.json (frozen before outcomes), evidence/summary-dev-val-run-v1.json, evidence/summary-v2-dev-val-run-v1.json (aggregates only; no symbols, dates or paths; each records its results hash, protocol hash and code revision).
  • Fees from primary SEC/FINRA/IBKR sources (fees.json). Private inputs stay off the repo; every page is ledger-hashed.
  • Elite data-rate use measured: the pre-market scan ran ~3,450 requests/min (34% of the 10,000/min limit).

Test plan

  • python3 scripts/validate.py passed
  • pinned runtime (Python 3.12.3, numpy 2.5.3, duckdb 1.5.5): 38 synthetic mover tests OK
  • python3 -m unittest (system Python): mover tests skip without numpy; test_effort_default_guard also fails on main on this host (it compares against the host-installed hook)
  • guarded gitleaks on every commit (standard limits): no leaks

🤖 Generated with Claude Code

…e data collection

protocol.json (mover-early-entry-v1-20260924), frozen before any post-entry return or quote
sample: long-only early-entry rules (8 times x 4 gain thresholds x 3 dollar-volume floors x 2
news variants x 4 exits), quote-measured costs from development sessions only, session-block
bootstrap with BH/Holm pass criteria over development/validation/holdout, capture by degree
tier, a preregistered 1x/2x/4x auto-leverage schedule with drawdown cooldown, capacity and
marginability caps, and data-quality gates. Revised from an adversarial pre-freeze review
(12 must-fix items adopted). candidates.py builds the provably complete superset (the day's
high >= 1.20 x the previous day's low; 236,538 symbol-days); collect.py and sessions_io.py
collect and read the private intraday data.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ions D1-D3 before any outcome

The frozen protocol (mover-early-entry-v1-20260924) is unchanged. This adds, before the cost table
or any post-entry return exists:

- rules.py: the protocol's definitions as pure functions (price and dollar volume at t, entry
  prices, exits X1-X4 with gap fills, halt flag, cost buckets, dated SEC/FINRA fees and the IBKR
  tiered commission, scalar and bitwise-equal vectorised net returns).
- features.py: per-symbol-day signals (pre-entry only) and outcomes, page hashes checked against
  each collection ledger, byte-identical output.
- quotes.py: the development-only 1-in-20 cost sample, SIP quote fetch and cost table.
- evaluate.py: metrics, circular session block bootstrap on quantised returns, BH/BY/Holm,
  two-way clustering, capture by degree tier, Wave H replication, coverage and completeness
  gates, and the leveraged 5-position portfolio (rungs 1/2/4, regime and drawdown factors).
- benchmarks.py (IWM official open-to-close), premarket_scan.py, mover_scan.py (live scanner that
  reuses rules.py so paper signals match the study).
- fees.json from primary SEC, FINRA and IBKR sources; clarifications.json (C1-C23).
- deviations.json: D1 an ordering slip (3-session exit trial deleted unread); D2 bars and
  auctions re-collected with asof 2026-09-21, the candidate list's naming date (asof = session
  missed 16.8% of development candidates); D3 daily bars are regular-session bars, so a
  pre-market scan adds 40,354 pre-market mover days the daily-high superset missed.

Synthetic tests: 22 (made-up symbols and prices).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…etups, pre-positioning) before any outcome

protocol-v2.json (mover-followup-v2-20260924) is frozen before any post-entry return of v1 or v2
exists. Families: E first-cross entries (216 rule-exits), F opening-range, pre-market-high and
VWAP-reclaim setups (16), P pre-positioning at the prior close and day-2 continuation (12).
Also adds the paper-versus-model fill comparison and a synthetic end-to-end pipeline test.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…y outcome

An independent review (3 Opus reviewers, 2 refuters per serious finding) of 1c61c78 found 10
blocking/major findings (7 distinct, all upheld) and 28 minor ones; all are resolved before any
post-entry return exists (deviations.json D5):
- the preregistered 600 s news sensitivity is reported per news rule-exit;
- portfolio slots rank every fire, filled or not (C24, no fill look-ahead);
- outcomes need the committed cost table and holdout reads need the dev_val results (C32);
- collections are refused on another asof or a symbol-count mismatch; ledger pages need files;
- the live scanner takes the previous session from the calendar, checks snapshot bar dates,
  waits for the signal cutoff, and prefilters today's splits; 09:30 rules are not paper
  candidates (C31);
- holdout statistics cover only the evaluated rule-exits; candidate rules C27-C30;
- P&L is booked on the cost basis; exact-threshold fires (C25); NYSE early closes (C26); the
  daily-close fallback (C22); the table's own 90th percentile (C9); coverage against the
  main list plus the D3 supplement; per-split completeness; input and runtime hash checks;
- D4 records the SPY history window; D3 records the post-dataset split gap (1 forward split).

evidence/cost-table-run-v1.json is the frozen cost table (aggregates only; sha256
be50cbdf...): 1,220 sampled development entries of 24,219 fired, 4,724 quote observations.
Also: v2 clarifications V1-V12. Synthetic tests: 27.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…evaluator before any v2 outcome

v1 (mover-early-entry-v1-20260924), dev_val on the frozen cost table, reproduced byte for byte by
two independent runs (results sha256 464f2bc2...): 0 of 768 rule-exits pass development; every
rule-exit with >= 200 development trades has a negative mean net return (median -2.9% per
trade); nothing is validated, so the holdout is not read. The paper candidate is the most-traded
non-09:30 development rule-exit, 10:00|G0.20|V250000|any|X1, labelled mechanics-only. Coverage of
the research package's verified +20% events is 83.8% (survivorship-limited label). Wave H:
62.8% of >= 100% gainers were +20% by 09:25 (package: 77%), Spearman 0.26 (package: 0.04).
evidence/summary-dev-val-run-v1.json holds the aggregates (no symbols or dates).

v2: features_v2.py (first-cross entries, opening-range / pre-market-high / VWAP-reclaim setups)
and evaluate_v2.py (families E, F, P with v1's statistics, costs and portfolio); clarifications
V13-V14; synthetic tests (3).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tudy in the evidence manifest

README: v1/v2 protocols, the v1 no-pass result, the reproduction steps for a new host and the file
roles. mover_scan.py can also write the adaptive engine's mover-trial scan file. The rules tests
skip without numpy like the other mover tests. manifests/evidence.json lists all study files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-24T05:39:12.509675Z f6790ac PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f6790ac25b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread blueprints/us-equities/mover-early-entry/features_v2.py Outdated
Comment thread blueprints/us-equities/mover-early-entry/evaluate_v2.py Outdated
Comment thread blueprints/us-equities/mover-early-entry/quotes.py Outdated
Comment thread blueprints/us-equities/mover-early-entry/evaluate_v2.py Outdated
Comment thread blueprints/us-equities/mover-early-entry/evaluate_v2.py Outdated
Comment thread blueprints/us-equities/mover-early-entry/paper_compare.py Outdated
Comment thread blueprints/us-equities/mover-early-entry/paper_compare.py Outdated
… any v2 outcome

v1 (post-outcome, labelled post hoc in D6): capture and replication also without basis-uncertain
rows (the C23 daily-close fallback or a split day) and with the C25 epsilon; selection and every
trade metric are unchanged (results 9d8c3d5b..., two runs byte-identical). The >= 10x tier holds
62 days, 16 without basis-uncertain rows; Wave H 62.9% / 69.7% (package: 77%). The README no
longer says validation is worse (it is similar) and states that -2.9% is the median of rule-exit
means. The public summary adds the protocol hash, the code revision and the remaining
preregistered per-rule-exit outputs.

v2 (independent review wf_a0e5858e-8a5: 9 serious findings, 8 upheld, plus 15 minor; all resolved
before any v2 outcome; clarifications V15-V20): P uses calendar-adjacent sessions, holds through
gaps of up to 5 sessions and books no row within 5 sessions as -100%, counts every skip, trades
each score's top 5 per session and books P&L after the next sizing; F3 enters after 09:31;
degenerate stops apply to Y2-Y4; every candidate day is written for E/F degree-tier capture;
BY, break-even, portfolio capacity runs and a holdout stage are added; the v1 cost table and
the v2 dev_val gate are pinned; seed provenance is fixed. Synthetic tests: 8 for v2.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… or P

evaluate_v2 dev_val on v1's frozen cost table, two independent runs byte-identical (results
sha256 3334962f...). E first-cross: all 216 rule-exits negative net (median -3.1%), none positive
before costs. F follow-up setups: all 16 negative net (best: 15-minute opening-range breakout with
its stop, -0.34%). P pre-positioning: all 12 negative net; the volume-breakout score's qualifying
set holds 11.7% of next-day >= +100% movers (base rate 0.004%) but loses on average after costs.
No holdout is read. evidence/summary-v2-dev-val-run-v1.json holds the aggregates.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins seathatflowsinourveins changed the title Mover early-entry studies: preregistered v1 (no pass), v2 follow-up protocol, data corrections D1-D5 Mover studies: preregistered v1 and v2 (no pass in any family), data corrections D1-D6 Sep 24, 2026
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Independent review of #162 by a separate Claude session (native-agent-stack-71). The review ran five lenses (preregistration, lookahead, costs, data, stats) with 100 agents, and two refuters checked every finding. It covered f6790ac and was re-checked against 31b7c2f. It does not cover 92062e8 (the v2 dev/val result) or the 11716ee merge. No market data was re-fetched. The tests pass (30/30 at f6790ac, 35/35 at 31b7c2f) only when numpy is installed. The system python3 skips all of them.

Verdict

  • The v1 "no pass" stands. The best development gross z is about 1.1 even at zero cost, so no change to the cost model could produce a pass.
  • v2's recorded "no pass" (92062e8) is probably robust but unreviewed. Items 3, 4 and 9 make the v2 cost model more lenient, so they bias toward passing, not against it. Items 1, 2 and 5-8 are integrity and reporting gaps. They should be closed with dated records before the v2 result is cited.
  • Item 2 matters most now. v2 had no pass, so the holdout must stay closed. At 31b7c2f a no-pass v2 dev_val could still open it.

Still open (medium)

  1. protocol-v2.json:7, README.md:9,13-14 (F6/S1/data-F5): the exposure text says everything was "fixed before the outcome". D6 and V13-V20 contradict that. The v2 code written after v1 dev/val was seen is not dated or listed. Fix: add a dated exposure deviation.
  2. features_v2.py:246-251, evaluate_v2.py:503-506 (F4): a v2 dev_val with no pass still opens the holdout. dev_val_sha256 is checked only when E/F have candidates. Fix: require a non-empty holdout_candidates and always check the hash.
  3. cost-table-run-v1.json:107-116, evaluate_v2.py:274-282 (costs-F1): the <$1M fallback costs about half as much as the $1-5M cell. E at $250k enters in that tier. Fix: make the fallback monotone or merge the two tiers.
  4. evaluate_v2.py:384-394 (costs-F6/S9/S7): P applies auction costs measured on movers to the whole universe. Its signal comes from the same close it enters at. A P pass needs no check against official prints, and P has no survivorship label. Fix: restrict P or re-measure its costs and add a pre-cutoff sensitivity, or label any P pass proxy-only.
  5. evaluate_v2.py:332-335,363,408 (LA-1): the S2/S3 windows count rows, not the sessions V8 specifies. p_capture drops gapped and delisted names from its base without counting them. Fix: use session windows and report the dropped count.
  6. features_v2.py:44-70,191 (LA-2, contested): E can fill at the 09:30 opening-print bar on a pre-market signal, and V16 covers only F3. Fix: add a C31-style 09:31 floor or next-bar entry, or flag those entries.
  7. features.py:277-280, evaluate.py:749-788 (F3): the v1 gate checks only stage and reads every holdout row. No access ledger backs "the holdout was not read". Fix: check the protocol id and a non-empty selection, and log holdout access.
  8. summarize.py:30-32 (F9): code_revision_at_summary is f6790ac, but f6790ac has no basis_uncertain code. The summary therefore names a revision that did not produce it. Fix: record the producing revision, code hashes and a receipt for each run.
  9. evaluate_v2.py:38 (costs-F4): pass statistics use a fixed $20k notional with no impact term and assume full fills. Fix: add an impact term or use the capacity-capped notional.
  10. protocol-v2.json:101-106 (S3): no minimum detectable effect is stated. Fix: state one per stage.
  11. candidates.py:5,43 (LA-3): the "provably complete" claim ignores the ~0.8% of days where the daily high and the minute highs disagree. Fix: measure which way they disagree, then add a tolerance or record a known gap.
  12. evaluate.py:621-638, README.md:22 (data-F1/LA-4): coverage does not resolve ticker renames, and the Wave H and capture claims leave out per-tier coverage. Fix: resolve tickers, then report coverage by cause and by tier.
  13. candidates.py:4-5, collect.py:3-4 (data-F2): the docstrings contradict D3, and D3 cites no commit.
  14. deviations.json:23-37, protocol-v2.json:88 (data-F3/F4): D2/D3 have no script or receipt, and the "78%" figure does not appear in D3.
  15. deviations.json:67-81 (data-F6): no review kept a findings receipt, and the v2 pre-outcome review has no id.
  16. paper_compare.py (costs-F5, contested): it models no per-side costs and covers v1 only.
  17. protocol-v2.json:102 (S5, contested, partly addressed): nothing corrects across v1 and v2 on the shared holdout. Fix: correct across them, or state the claims per family.

Low: mover_scan.py:180-184 drops names halted for the whole previous day. candidates.py:43 and premarket_scan.py:153 have no epsilon. Auction sides are charged 1.25x vs 0.5x. Stops fill at their level. Fees have no source. Pre-market halts are not flagged. D2/D3 have no amended_at, no UNX rescan and no record of which way the mismatches go. The README repro file names are wrong. CI runs no mover tests.

Addressed by 31b7c2f

prereg-F1, F7, F8, F10; LA-7; costs-F2; S2, S6, S8. V17 changes V7's skip rule, so it should be recorded as a deviation.

…exact quote reuse, paper comparison identity

- P exits use the exit row's raw price, split-adjusted shares and dividend cash (D7, post-outcome,
  labelled): E/F identical, P means change by <= 0.00002, still no pass (results 75a5004b...).
- quotes.py: completions record their timestamp and are reused only for the same one; table
  refuses quotes outside a stamp's window (the frozen cost table reproduces byte for byte).
- paper_compare.py: asof pinned to the trade session; the daily-close fallback as the study.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins merged commit aa6fc79 into main Sep 24, 2026
22 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the claude/mover-early-entry-20260924 branch September 24, 2026 06:41
seathatflowsinourveins pushed a commit that referenced this pull request Sep 24, 2026
…nd catalog decisions

blueprints/us-equities/mover-v3/README.md (2026-09-24) turns the verified sweep
(37 items: 5 survive, 24 contested, 8 killed) into a plan. It uses only the
surviving items and the contested items whose refutation the synthesis judged
weak, labelled as contested. It covers why continuation is not supported
(BCW 2011 MAX, Hong et al. days to cover, the Barber/Odean prior, #162's
0/768 and v2 no-pass), H1-H6 with entry, exit, sizing, falsifier and data
status, the data plan (available now vs Databento/Norgate/FINRA gated), the
broker-held exit and leverage-ladder policy, the call-cost table for the
200/min and 1000/min tiers (throughput is capacity, not a target; the plain
limit pattern is corrected to 2/f calls per round trip), simulation fidelity,
and the dropped items with reasons.

protocol-draft.json is status draft_pending_independent_pre_outcome_review,
not frozen. It sets windows disjoint from v1/v2 (development 2017-2019,
validation 2020, a prospective holdout) and names the prior exposures. One
Holm family of 37 items at 0.05 covers H1-H6, with a lineage sensitivity
over 1,012 earlier rule-exits and MDE tables. The cost model reuses #162's v1
table by sha256 be50cbdf, makes it monotone in dollar volume (7 cells) and
adds capacity caps plus an impact term. The holdout gate has a committed
access log.

catalogs/us-equities/mover-v3-sweep-20260924.json holds the add, compare and
no-change records with their evidence levels. Its /repository_decisions (7
existing repositories) is registered via catalog_decisions.py --write
--supplement; the repository count is unchanged (844), and there are 1,968
references. manifests/evidence.json updates the reviewed hashes of the two
changed registered files. PR #162's files are untouched. This is literature
and metadata evidence only: no native execution, no data fetched, and no
outcome computed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 24, 2026
…tions, repin to merged #162

Rebased on main aa6fc79 (PR #162 merged) and repinned v1 reuse to main with
blob and sha256 hashes for rules.py, quotes.py (with the D7 stamp-window
guard) and the unchanged cost table; cite D7 and the v2 results sha 75a5004b.

Protocol draft: record broad-universe's $1/$2M descriptive lane and the S2
headline as aggregate exposure of 2017-2020, so development and validation
are screening only and only the prospective holdout is confirmatory. Move H6
to its own engineering draft and two-test family (m = 35 alpha family), with
TOST-sized stages, a TOST p for Holm and a hash key the client cannot re-roll.
Trailing 252-session tercile breakpoints; three preregistered outcome labels;
leg-matched H3-c on official auction prints; no halt data in H1-H3; a single
deduplicated H4 parent set with a scoped holdout read; B = 100,000 so the
lineage level is reachable; two-sided MDEs; pinned cost lookup inputs, sample
key and sparse-cell fallbacks; holdout starts after the validation results are
committed, with accrual exposure logged; embargo side stated.

Plan and catalog: one broker-held closing order per share (Alpaca docs quoted),
TotalView-ITCH coverage and frames accn corrected, squeeze claim detached from
Hong et al., overnight/intraday primaries with Crossref-resolved DOIs, contested
labels on Databento and Norgate, backtrader registered and the Qullamaggie
scanner left unregistered with its reason; evidence hashes for the new files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 24, 2026
…gate H2 and H6

Resolve the 27 findings of the two rereviews of 0b924e7 in the draft
protocol, the H6 engineering draft and the README. Status stays
draft_pending_independent_pre_outcome_review; frozen_before_outcomes false.

- H3-c and H4 are two-sided only at development; the development sign
  locks a one-sided test at validation and the holdout.
- H2, H4 and H5 are recorded at the freeze as tested or permanently
  not tested; later data needs a new protocol version.
- The 2017-2019 cost sample moves after the freeze (its stamps are
  outcome prices).
- Defined MDE sigmas, label bound, unequal groups, H3-c leg rule,
  holdout breakpoints outside #162's window, fixed holdout start,
  count-only access, segment-end censoring, early-close times, mid
  fills, D7 split/dividend accounting (evaluate_v2.py pinned), lineage
  denominator 1,107, tie-breaks and PCG64 streams.
- H2 moved to gated: catalyst.py fails every historical filing's
  availability and has no Item parser or historical CIK map.
- H6 gated with no position source; H6-c withdrawn as a test because
  the client-side arm needs fewer calls by construction.
- Evidence manifest hashes recomputed for the three changed files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 24, 2026
…h merged evidence (#170)

Records only, no gate status change: runtime-target IBKR cites the passed 1.231.0 paper receipt (#147) while rc5 local acceptance stays not_established (blocker nautilus#4983); dashboard checkpoint cites the 32-layer sweep (#153) and the #162 mover research result. Independently reviewed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 24, 2026
…rified; paper needs an isolated account) (#172)

Bounded paper trial mode for the mover studies (#162): scanner-selected symbols (at most 5), marketable-limit entries (extended hours outside RTH), the preregistered leverage schedule with capacity and sub-$5 caps, one study exit rule (X1-X4, X1 two minutes before the session's regular close), a hard flatten, reconciliation and a receipt that keeps the order journal. New modules only; shared engine modules unchanged. Two independent reviews plus rechecks (2 blocking and several major findings fixed with failing-first tests) and a Codex review (8 threads resolved). Evidence: SYN through the real LiveNode on the live scanner's output (5 entries, 5 X2 exits, flat, reconciled, cash delta = realized P&L); 732 adaptive-paper tests OK on both runtimes. No PAPER evidence yet: the mover lane needs its own Alpaca paper account (it must not share the adaptive lane's).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 24, 2026
…mance, and the standing live-readiness instruction

Two required live gates, criteria fixed before any strategy qualifies and
revised before merge after an independent review.

strategy-out-of-sample-holdout is established by ONE of two routes: (i) the
strategy's frozen protocol passes its own holdout rule, with its minimum trades
and multiple-testing adjustment and any lower bound at the family-wise adjusted
level (mover-early-entry v1: <= 5 rule-exits, >= 50 trades each, Holm p <= 0.05,
paper candidate per paper_e2e); or (ii) a preregistered prospective paper study,
pinned by sha256 before its first order, passes its own test at its single
analysis point. Both need >= 100 round trips and 20 sessions (or the protocol's
larger minimums), costs from a sha256-pinned cost table (per-side half-spread
plus the protocol's slippage rule) and fee schedule, and recorded market data or
real Alpaca paper fills pinned by sha256; synthetic, offline or fixture evidence
never counts.

strategy-paper-performance needs a separate, later paper period with the same
frozen configs (no session counts toward both gates), one preregistered analysis
point (first session end with >= 20 sessions and >= 100 round trips, analysed
once), a pooled per-round-trip net estimator with a one-sided 95% circular
5-session block bootstrap lower bound > 0, drawdown <= 20%, zero needs_attention,
per-side adverse slippage (entry fill/model - 1, exit model/fill - 1) with a
median within the pinned cost assumption via the strategy's own comparator,
scheduled trading days as sessions, and leverage above 1x only at an established
ladder rung.

Both gates share blueprints/us-equities/strategy-performance/receipt.json; its
sections must carry the same strategy_id, protocol_sha256 and config_sha256, and
the checker verifies only the status strings. The live-go note records the
user's 2026-09-24 instruction: paper and simulation are not subject to human
approval steps; the live lane is ready for the user's decision when the checker's
blocking.live equals ["live-go"]; the user stores the live key and gives the
explicit go; automation never flips live-go or places live orders. The handbook
table is dated September 24 (22 gates, 21 checked) and its paper row matches the
checker. Tests add both gates to the GOOD/BAD receipt fixtures and check that an
out_of_sample-only receipt flips only that gate. No strategy qualifies today
(#162).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant