Repository navigation
Mover studies: preregistered v1 and v2 (no pass in any family), data corrections D1-D6 - #162
Conversation
…e data collection protocol.json (mover-early-entry-v1-20260924), frozen before any post-entry return or quote sample: long-only early-entry rules (8 times x 4 gain thresholds x 3 dollar-volume floors x 2 news variants x 4 exits), quote-measured costs from development sessions only, session-block bootstrap with BH/Holm pass criteria over development/validation/holdout, capture by degree tier, a preregistered 1x/2x/4x auto-leverage schedule with drawdown cooldown, capacity and marginability caps, and data-quality gates. Revised from an adversarial pre-freeze review (12 must-fix items adopted). candidates.py builds the provably complete superset (the day's high >= 1.20 x the previous day's low; 236,538 symbol-days); collect.py and sessions_io.py collect and read the private intraday data. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ions D1-D3 before any outcome The frozen protocol (mover-early-entry-v1-20260924) is unchanged. This adds, before the cost table or any post-entry return exists: - rules.py: the protocol's definitions as pure functions (price and dollar volume at t, entry prices, exits X1-X4 with gap fills, halt flag, cost buckets, dated SEC/FINRA fees and the IBKR tiered commission, scalar and bitwise-equal vectorised net returns). - features.py: per-symbol-day signals (pre-entry only) and outcomes, page hashes checked against each collection ledger, byte-identical output. - quotes.py: the development-only 1-in-20 cost sample, SIP quote fetch and cost table. - evaluate.py: metrics, circular session block bootstrap on quantised returns, BH/BY/Holm, two-way clustering, capture by degree tier, Wave H replication, coverage and completeness gates, and the leveraged 5-position portfolio (rungs 1/2/4, regime and drawdown factors). - benchmarks.py (IWM official open-to-close), premarket_scan.py, mover_scan.py (live scanner that reuses rules.py so paper signals match the study). - fees.json from primary SEC, FINRA and IBKR sources; clarifications.json (C1-C23). - deviations.json: D1 an ordering slip (3-session exit trial deleted unread); D2 bars and auctions re-collected with asof 2026-09-21, the candidate list's naming date (asof = session missed 16.8% of development candidates); D3 daily bars are regular-session bars, so a pre-market scan adds 40,354 pre-market mover days the daily-high superset missed. Synthetic tests: 22 (made-up symbols and prices). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…etups, pre-positioning) before any outcome protocol-v2.json (mover-followup-v2-20260924) is frozen before any post-entry return of v1 or v2 exists. Families: E first-cross entries (216 rule-exits), F opening-range, pre-market-high and VWAP-reclaim setups (16), P pre-positioning at the prior close and day-2 continuation (12). Also adds the paper-versus-model fill comparison and a synthetic end-to-end pipeline test. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…y outcome An independent review (3 Opus reviewers, 2 refuters per serious finding) of 1c61c78 found 10 blocking/major findings (7 distinct, all upheld) and 28 minor ones; all are resolved before any post-entry return exists (deviations.json D5): - the preregistered 600 s news sensitivity is reported per news rule-exit; - portfolio slots rank every fire, filled or not (C24, no fill look-ahead); - outcomes need the committed cost table and holdout reads need the dev_val results (C32); - collections are refused on another asof or a symbol-count mismatch; ledger pages need files; - the live scanner takes the previous session from the calendar, checks snapshot bar dates, waits for the signal cutoff, and prefilters today's splits; 09:30 rules are not paper candidates (C31); - holdout statistics cover only the evaluated rule-exits; candidate rules C27-C30; - P&L is booked on the cost basis; exact-threshold fires (C25); NYSE early closes (C26); the daily-close fallback (C22); the table's own 90th percentile (C9); coverage against the main list plus the D3 supplement; per-split completeness; input and runtime hash checks; - D4 records the SPY history window; D3 records the post-dataset split gap (1 forward split). evidence/cost-table-run-v1.json is the frozen cost table (aggregates only; sha256 be50cbdf...): 1,220 sampled development entries of 24,219 fired, 4,724 quote observations. Also: v2 clarifications V1-V12. Synthetic tests: 27. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…evaluator before any v2 outcome v1 (mover-early-entry-v1-20260924), dev_val on the frozen cost table, reproduced byte for byte by two independent runs (results sha256 464f2bc2...): 0 of 768 rule-exits pass development; every rule-exit with >= 200 development trades has a negative mean net return (median -2.9% per trade); nothing is validated, so the holdout is not read. The paper candidate is the most-traded non-09:30 development rule-exit, 10:00|G0.20|V250000|any|X1, labelled mechanics-only. Coverage of the research package's verified +20% events is 83.8% (survivorship-limited label). Wave H: 62.8% of >= 100% gainers were +20% by 09:25 (package: 77%), Spearman 0.26 (package: 0.04). evidence/summary-dev-val-run-v1.json holds the aggregates (no symbols or dates). v2: features_v2.py (first-cross entries, opening-range / pre-market-high / VWAP-reclaim setups) and evaluate_v2.py (families E, F, P with v1's statistics, costs and portfolio); clarifications V13-V14; synthetic tests (3). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tudy in the evidence manifest README: v1/v2 protocols, the v1 no-pass result, the reproduction steps for a new host and the file roles. mover_scan.py can also write the adaptive engine's mover-trial scan file. The rules tests skip without numpy like the other mover tests. manifests/evidence.json lists all study files. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f6790ac25b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
… any v2 outcome v1 (post-outcome, labelled post hoc in D6): capture and replication also without basis-uncertain rows (the C23 daily-close fallback or a split day) and with the C25 epsilon; selection and every trade metric are unchanged (results 9d8c3d5b..., two runs byte-identical). The >= 10x tier holds 62 days, 16 without basis-uncertain rows; Wave H 62.9% / 69.7% (package: 77%). The README no longer says validation is worse (it is similar) and states that -2.9% is the median of rule-exit means. The public summary adds the protocol hash, the code revision and the remaining preregistered per-rule-exit outputs. v2 (independent review wf_a0e5858e-8a5: 9 serious findings, 8 upheld, plus 15 minor; all resolved before any v2 outcome; clarifications V15-V20): P uses calendar-adjacent sessions, holds through gaps of up to 5 sessions and books no row within 5 sessions as -100%, counts every skip, trades each score's top 5 per session and books P&L after the next sizing; F3 enters after 09:31; degenerate stops apply to Y2-Y4; every candidate day is written for E/F degree-tier capture; BY, break-even, portfolio capacity runs and a holdout stage are added; the v1 cost table and the v2 dev_val gate are pinned; seed provenance is fixed. Synthetic tests: 8 for v2. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… or P evaluate_v2 dev_val on v1's frozen cost table, two independent runs byte-identical (results sha256 3334962f...). E first-cross: all 216 rule-exits negative net (median -3.1%), none positive before costs. F follow-up setups: all 16 negative net (best: 15-minute opening-range breakout with its stop, -0.34%). P pre-positioning: all 12 negative net; the volume-breakout score's qualifying set holds 11.7% of next-day >= +100% movers (base rate 0.004%) but loses on average after costs. No holdout is read. evidence/summary-v2-dev-val-run-v1.json holds the aggregates. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Independent review of #162 by a separate Claude session (native-agent-stack-71). The review ran five lenses (preregistration, lookahead, costs, data, stats) with 100 agents, and two refuters checked every finding. It covered f6790ac and was re-checked against 31b7c2f. It does not cover 92062e8 (the v2 dev/val result) or the 11716ee merge. No market data was re-fetched. The tests pass (30/30 at f6790ac, 35/35 at 31b7c2f) only when numpy is installed. The system python3 skips all of them. Verdict
Still open (medium)
Low: Addressed by 31b7c2fprereg-F1, F7, F8, F10; LA-7; costs-F2; S2, S6, S8. V17 changes V7's skip rule, so it should be recorded as a deviation. |
…exact quote reuse, paper comparison identity - P exits use the exit row's raw price, split-adjusted shares and dividend cash (D7, post-outcome, labelled): E/F identical, P means change by <= 0.00002, still no pass (results 75a5004b...). - quotes.py: completions record their timestamp and are reused only for the same one; table refuses quotes outside a stamp's window (the frozen cost table reproduces byte for byte). - paper_compare.py: asof pinned to the trade session; the daily-close fallback as the study. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…nd catalog decisions blueprints/us-equities/mover-v3/README.md (2026-09-24) turns the verified sweep (37 items: 5 survive, 24 contested, 8 killed) into a plan. It uses only the surviving items and the contested items whose refutation the synthesis judged weak, labelled as contested. It covers why continuation is not supported (BCW 2011 MAX, Hong et al. days to cover, the Barber/Odean prior, #162's 0/768 and v2 no-pass), H1-H6 with entry, exit, sizing, falsifier and data status, the data plan (available now vs Databento/Norgate/FINRA gated), the broker-held exit and leverage-ladder policy, the call-cost table for the 200/min and 1000/min tiers (throughput is capacity, not a target; the plain limit pattern is corrected to 2/f calls per round trip), simulation fidelity, and the dropped items with reasons. protocol-draft.json is status draft_pending_independent_pre_outcome_review, not frozen. It sets windows disjoint from v1/v2 (development 2017-2019, validation 2020, a prospective holdout) and names the prior exposures. One Holm family of 37 items at 0.05 covers H1-H6, with a lineage sensitivity over 1,012 earlier rule-exits and MDE tables. The cost model reuses #162's v1 table by sha256 be50cbdf, makes it monotone in dollar volume (7 cells) and adds capacity caps plus an impact term. The holdout gate has a committed access log. catalogs/us-equities/mover-v3-sweep-20260924.json holds the add, compare and no-change records with their evidence levels. Its /repository_decisions (7 existing repositories) is registered via catalog_decisions.py --write --supplement; the repository count is unchanged (844), and there are 1,968 references. manifests/evidence.json updates the reviewed hashes of the two changed registered files. PR #162's files are untouched. This is literature and metadata evidence only: no native execution, no data fetched, and no outcome computed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tions, repin to merged #162 Rebased on main aa6fc79 (PR #162 merged) and repinned v1 reuse to main with blob and sha256 hashes for rules.py, quotes.py (with the D7 stamp-window guard) and the unchanged cost table; cite D7 and the v2 results sha 75a5004b. Protocol draft: record broad-universe's $1/$2M descriptive lane and the S2 headline as aggregate exposure of 2017-2020, so development and validation are screening only and only the prospective holdout is confirmatory. Move H6 to its own engineering draft and two-test family (m = 35 alpha family), with TOST-sized stages, a TOST p for Holm and a hash key the client cannot re-roll. Trailing 252-session tercile breakpoints; three preregistered outcome labels; leg-matched H3-c on official auction prints; no halt data in H1-H3; a single deduplicated H4 parent set with a scoped holdout read; B = 100,000 so the lineage level is reachable; two-sided MDEs; pinned cost lookup inputs, sample key and sparse-cell fallbacks; holdout starts after the validation results are committed, with accrual exposure logged; embargo side stated. Plan and catalog: one broker-held closing order per share (Alpaca docs quoted), TotalView-ITCH coverage and frames accn corrected, squeeze claim detached from Hong et al., overnight/intraday primaries with Crossref-resolved DOIs, contested labels on Databento and Norgate, backtrader registered and the Qullamaggie scanner left unregistered with its reason; evidence hashes for the new files. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…gate H2 and H6 Resolve the 27 findings of the two rereviews of 0b924e7 in the draft protocol, the H6 engineering draft and the README. Status stays draft_pending_independent_pre_outcome_review; frozen_before_outcomes false. - H3-c and H4 are two-sided only at development; the development sign locks a one-sided test at validation and the holdout. - H2, H4 and H5 are recorded at the freeze as tested or permanently not tested; later data needs a new protocol version. - The 2017-2019 cost sample moves after the freeze (its stamps are outcome prices). - Defined MDE sigmas, label bound, unequal groups, H3-c leg rule, holdout breakpoints outside #162's window, fixed holdout start, count-only access, segment-end censoring, early-close times, mid fills, D7 split/dividend accounting (evaluate_v2.py pinned), lineage denominator 1,107, tie-breaks and PCG64 streams. - H2 moved to gated: catalyst.py fails every historical filing's availability and has no Item parser or historical CIK map. - H6 gated with no position source; H6-c withdrawn as a test because the client-side arm needs fewer calls by construction. - Evidence manifest hashes recomputed for the three changed files. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…h merged evidence (#170) Records only, no gate status change: runtime-target IBKR cites the passed 1.231.0 paper receipt (#147) while rc5 local acceptance stays not_established (blocker nautilus#4983); dashboard checkpoint cites the 32-layer sweep (#153) and the #162 mover research result. Independently reviewed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rified; paper needs an isolated account) (#172) Bounded paper trial mode for the mover studies (#162): scanner-selected symbols (at most 5), marketable-limit entries (extended hours outside RTH), the preregistered leverage schedule with capacity and sub-$5 caps, one study exit rule (X1-X4, X1 two minutes before the session's regular close), a hard flatten, reconciliation and a receipt that keeps the order journal. New modules only; shared engine modules unchanged. Two independent reviews plus rechecks (2 blocking and several major findings fixed with failing-first tests) and a Codex review (8 threads resolved). Evidence: SYN through the real LiveNode on the live scanner's output (5 entries, 5 X2 exits, flat, reconciled, cash delta = realized P&L); 732 adaptive-paper tests OK on both runtimes. No PAPER evidence yet: the mover lane needs its own Alpaca paper account (it must not share the adaptive lane's). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…mance, and the standing live-readiness instruction Two required live gates, criteria fixed before any strategy qualifies and revised before merge after an independent review. strategy-out-of-sample-holdout is established by ONE of two routes: (i) the strategy's frozen protocol passes its own holdout rule, with its minimum trades and multiple-testing adjustment and any lower bound at the family-wise adjusted level (mover-early-entry v1: <= 5 rule-exits, >= 50 trades each, Holm p <= 0.05, paper candidate per paper_e2e); or (ii) a preregistered prospective paper study, pinned by sha256 before its first order, passes its own test at its single analysis point. Both need >= 100 round trips and 20 sessions (or the protocol's larger minimums), costs from a sha256-pinned cost table (per-side half-spread plus the protocol's slippage rule) and fee schedule, and recorded market data or real Alpaca paper fills pinned by sha256; synthetic, offline or fixture evidence never counts. strategy-paper-performance needs a separate, later paper period with the same frozen configs (no session counts toward both gates), one preregistered analysis point (first session end with >= 20 sessions and >= 100 round trips, analysed once), a pooled per-round-trip net estimator with a one-sided 95% circular 5-session block bootstrap lower bound > 0, drawdown <= 20%, zero needs_attention, per-side adverse slippage (entry fill/model - 1, exit model/fill - 1) with a median within the pinned cost assumption via the strategy's own comparator, scheduled trading days as sessions, and leverage above 1x only at an established ladder rung. Both gates share blueprints/us-equities/strategy-performance/receipt.json; its sections must carry the same strategy_id, protocol_sha256 and config_sha256, and the checker verifies only the status strings. The live-go note records the user's 2026-09-24 instruction: paper and simulation are not subject to human approval steps; the live lane is ready for the user's decision when the checker's blocking.live equals ["live-go"]; the user stores the live key and gives the explicit go; automation never flips live-go or places live orders. The handbook table is dated September 24 (22 gates, 21 checked) and its paper row matches the checker. Tests add both gates to the GOOD/BAD receipt fixtures and check that an out_of_sample-only receipt flips only that gate. No strategy qualifies today (#162). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Summary
Preregistered, reproducible tests of whether long-only rules can pre-position in or enter US-equity extreme movers early, net of quote-measured costs, with a preregistered auto-leverage schedule (rungs 1/2/4 x regime x drawdown factors, cooldown, capacity caps).
protocol.json, frozen4e57667a): 8 clock times x thresholds x volume floors x news x 4 exits = 768 rule-exits. No development pass; all 704 rule-exits with >= 200 trades are negative net (median of rule-exit means -2.9%); validation similar (-2.7%). Holdout not read. Results reproduced byte for byte (sha2569d8c3d5b...).protocol-v2.json, frozena4a2f682before any v1 or v2 outcome): E first-cross (earliest-stage) entries (216), F opening-range / pre-market-high / VWAP-reclaim setups (16), P pre-positioning at the prior close and day-2 continuation (12). No development pass in any family: E all negative even before costs; F best -0.34% net; P best -1.0% net, though its volume-breakout score holds 11.7% of next-day >= +100% movers (base rate 0.004%). Results sha2563334962f..., byte-identical reruns.deviations.json): D2 bars and auctions re-collected with the candidate list's naming date (asof = session lost 16.8% of candidates); D3 Alpaca daily bars are regular-session bars, so a pre-market scan (36,185 pages, all 200) added 40,354 pre-market mover days; D4 SPY window; D1 an ordering slip (deleted unread); D5 the independent v1 pre-outcome review (10 serious findings, resolved). D6 (post-outcome, labelled): degree-tier capture and replication also without basis-uncertain rows; trade metrics and selection unchanged.mover_scan.py, full market in ~6 s, 33 requests) reuses the study's rules;paper_compare.pyscores paper fills against the historical fill model.Evidence
evidence/cost-table-run-v1.json(frozen before outcomes),evidence/summary-dev-val-run-v1.json,evidence/summary-v2-dev-val-run-v1.json(aggregates only; no symbols, dates or paths; each records its results hash, protocol hash and code revision).fees.json). Private inputs stay off the repo; every page is ledger-hashed.Test plan
python3 scripts/validate.pypassedpython3 -m unittest(system Python): mover tests skip without numpy;test_effort_default_guardalso fails on main on this host (it compares against the host-installed hook)🤖 Generated with Claude Code