Repository navigation
Definitive round: a clean-room, two-family audit of the definitive manifest on 43 slots, with deep-dive dossiers - #663
seathatflowsinourveins wants to merge 9 commits into
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
77f9e83 to
08227eb
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 77f9e836ad
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| @@ -209,6 +209,12 @@ | |||
| "tests/test_catalogs.py", | |||
| "tests/test_new_host_grand_list.py", | |||
| "tests/test_handbook_summary.py", | |||
| # The clean-room definitive round of 2026-10-02 (records that name each slot's picks, the manifest rows they map | |||
| # to and the comparison with the manifest): every top-level file of its folder, its decision record and the | |||
There was a problem hiding this comment.
Withhold all top-level round files from blind exports
When a later lane uses blind_checkout.py at this revision, this glob removes only top-level JSON files despite the comment promising every top-level file. It leaves assemble.py, whose USER_PINS directly names the selected OmniRoute, OpenHands, and GPT Researcher repositories, and build_units.py, whose source-slot keys associate prior candidates with jobs; a supposedly blind reviewer can therefore recover prior selections and contaminate the experiment. Remove all non-dossier top-level round files and test representative .py/.md files.
Useful? React with 👍 / 👎.
| if isinstance(record, dict) and "attempts" in record: | ||
| attempts.append({"file": rel, "attempts": record["attempts"]}) |
There was a problem hiding this comment.
Record dossier attempts in the assembled run record
When this assembler processes dossier records, they store attempts under writer and nested verification[], not a top-level attempts key, so this condition silently excludes every dossier writer and verifier. The committed run-record.json consequently contains only the 150 decision/critic/adjudication groups and omits the attempts behind all 177 dossiers, although the README says it records every attempt; this loses their exit, duration, and usage evidence and makes the published accounting incomplete. Collect the dossier-specific attempt fields as well.
AGENTS.md reference: AGENTS.md:L13-L14
Useful? React with 👍 / 👎.
| for _ in range(2): | ||
| r = fn(*args, **kwargs) | ||
| attempts.append({k: r[k] for k in ("exit", "seconds", "usage", "stderr_tail")}) |
There was a problem hiding this comment.
Preserve each retry's raw return separately
Whenever a model call needs the second attempt, both invocations receive the same raw path, so run() overwrites the first attempt's stdout and stderr with the retry. The committed run record contains three two-attempt jobs, meaning their first raw event streams are no longer among the hashed private originals even though the experiment claims to preserve every attempt; use an attempt-indexed raw path before retrying.
AGENTS.md reference: AGENTS.md:L18-L18
Useful? React with 👍 / 👎.
| def gpt(prompt: str, cwd: Path, schema_path: Path, timeout: int, raw: Path, work: Path) -> dict: | ||
| env = dict(os.environ, CODEX_HOME=str(gpt_home(work)), OMNIROUTE_API_KEY="local-loopback") | ||
| last = raw.with_suffix(".last.json") | ||
| cmd = ["codex", "exec", "--skip-git-repo-check", "--ephemeral", "-s", "read-only", "-C", str(cwd), | ||
| "-m", GPT_MODEL, "-c", f'model_reasoning_effort="{GPT_EFFORT}"', "-c", 'web_search="live"', "--json", | ||
| "--output-schema", str(schema_path), "-o", str(last), prompt] | ||
| result = run(cmd, cwd, timeout, env=env, stdout_path=raw) |
There was a problem hiding this comment.
Create the GPT output directory before invoking Codex
On the first GPT call in each fresh stage, last is inside the not-yet-created raw parent directory, while run() creates that directory only after the subprocess exits. Codex therefore cannot write -o, returns no parsed result, and triggers a full expensive retry; the committed record shows this exact Failed to write last message file ... No such file or directory failure for three calls. Create last.parent before starting Codex.
Useful? React with 👍 / 👎.
| def stage_decide(work: Path, jobs: int, only: set, families: list) -> None: | ||
| tasks = [(fam, u, o) for fam in families for u in units() if not only or u["unit_id"] in only for o in (1, 2)] | ||
| with ThreadPoolExecutor(max_workers=jobs) as pool: | ||
| results = list(pool.map(lambda t: decide_one(t, work), tasks)) |
There was a problem hiding this comment.
Run the cross-family A/B through promptfoo
This stage implements a bespoke LLM A/B runner over both model families and both candidate orders, but the repository explicitly requires gateway and LLM A/B work to use promptfoo and forbids a self-written runner. Because the definitive outcomes and their evidence were produced by this custom orchestration rather than the required upstream harness, the round does not satisfy the repository's acceptance process and should be rerun through promptfoo with the same frozen inputs.
AGENTS.md reference: AGENTS.md:L3-L3
Useful? React with 👍 / 👎.
| if len(adjudications) == 4 and len(resolved) == 4 and len(set(resolved)) == 1 and resolved[0] is not None: | ||
| result = {"status": "definitive", "basis": "adjudicated (four of four)", "default": resolved[0]} | ||
| else: | ||
| finalists = sorted({x for x in (c, g) if x}) |
There was a problem hiding this comment.
Preserve the second finalist when one family is split
For an unsettled slot where one critic is split, this expression builds finalists only from the families' surviving picks and therefore drops every candidate represented by the split return. This already corrupts the published document-ingestion result: selection.json lists only F07 as a finalist even though its recorded settling measurement explicitly compares F05 with F07. Preserve explicit finalists from split critics/adjudication rather than deriving them solely from non-null family picks.
Useful? React with 👍 / 👎.
| measurement = next((a["result"].get("settling_measurement") for a in adjudications | ||
| if a.get("result") and a["result"].get("settling_measurement")), None) |
There was a problem hiding this comment.
Supply a settlement measure for every measurement outcome
If every adjudicator leaves settling_measurement null, this code still emits a measurement outcome with no way to settle it and never falls back to the unit's preregistered deciding comparison. That occurs in the committed results for local-model-server, workflow-engine, and page-text-extraction, leaving three supposedly measurement-bound outcomes without an actionable measurement. Use the unit-level comparison or critic measurements as a fallback, or keep the result explicitly unresolved rather than labeling it as a measurement.
Useful? React with 👍 / 👎.
| "schema_version": 1, | ||
| "kind": "definitive_round_run_record", | ||
| "attempts": [ |
There was a problem hiding this comment.
Add a validator-compatible convergence experiment record
For this new convergence audit, the committed run record is not a scoped experiment record accepted by the repository's required validator: python3 scripts/validate_convergence.py evidence/artifacts/new-wsl-definitive-round-20261002/run-record.json --root . --json reports missing required fields, unknown fields, and an unsupported kind, and no other validator-compatible record for the round was added. The convergence claims therefore bypass the mandated artifact-hash, condition, failure-retention, and qualification consistency checks; add and validate the required contract/receipt before treating this as a completed round.
AGENTS.md reference: AGENTS.md:L18-L18
Useful? React with 👍 / 👎.
| FROZEN = ["criteria.txt", "units.json", "contenders.json", "build_units.py", "dossier-prompt.txt", "dossier-schema.json", | ||
| "verify-prompt.txt", "verify-schema.json", "decide-prompt.txt", "decide-schema.json", "critic-prompt.txt", | ||
| "critic-schema.json", "adjudicate-prompt.txt", "adjudicate-schema.json", "decision-rule.txt", | ||
| "clean-room.json", "run_round.py"] |
There was a problem hiding this comment.
Freeze the assembler that determines the round outcome
The preregistration omits assemble.py from FROZEN, even though that file defines the user pins and the exact logic that converts critic/adjudicator returns into definitive, measurement, and conflict outcomes. A post-result edit to that scoring implementation can therefore regenerate selection.json while freeze.py --check still reports unchanged, without any recorded amendment. Include the assembler's hash in the initial preregistration or amendment chain; the current check cannot establish that the published outcome used the preregistered implementation.
Useful? React with 👍 / 👎.
|
FINDINGS — exact head
Bounded SOURCE ACCEPT: all215 owned byte/SHA/Git identities and212registered bindings reconcile, with191 complete-added-file reconstructions from original immutable diff. Against frozen4ced:9631file rows preserve9420base membership/order and9419foreign rows;187receipts/26convergence records unchanged. Category counts reconcile22agree/sevencontest/sevennot-settled/onepin-conflict/onepin-agree/onecross-check/fournominations. The21freeze/amendment inputs bind; recovered pre-amendment4 source proves the disclosed20/9→22/7 post-result change, with pre-fix report retained. It is not result-neutral. No definitive-manifest/default row changed. Original source-capture failures remain32native0/threeREST403native1 plus the separate processingKeyError1; two native GraphQL recoveries supplement them. The Astra reader's19 source commands all exited0; it executed no tests/repository helper, benchmark, provider/GPU qualification or installation. Terminal CLI-scope only usage, counted once: input1614185/cached-input1462144/cache-write0/output27954/reasoning-output13304; subsets are not added, served identity and provider-complete accounting remain unknown. Private originals, full original comparison-input bodies and historical intermediate runner bodies were outside the frozen packet, so related semantics/execution remain unqualified. Current59f8 carry/hosted gates, all comparative role/default admissions and comparison/freeze HOLD remain separate. Repairs belong to the publication owner; root edits no Claude-lane source. |
19 units and 43 open foundation slots of the new-WSL definitive defaults, each asked as a neutral function question, with every candidate of each unit's field (177 contenders: the frozen #589 packets plus the runtime rows' fields). Every candidate gets a dossier from its upstream source at its release tag (README read in full, code, packaging, tests and CI, README claims checked against code), verified once and repaired once on a failed check. Both families then decide per unit in two orders, a critic per family checks both deciders, and contested slots go to adjudicators of both families in both A/B orders; a slot adjudication does not settle keeps two finalists and the named measurement. Every model call runs in a clean room: claude -p in safe and restricted mode, and codex exec with an isolated CODEX_HOME on the OmniRoute gateway; clean-room.json records the probes. preregistration.json holds the sha256 of every frozen file, recorded at 2026-10-02T03:04:22Z before the first dossier. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A smoke decision on a toy unit outside the field showed that the Claude CLI's multi-value --add-dir and --allowedTools flags read the prompt as one more value once clone directories are added, so every Claude decider, critic and adjudicator would have exited without running. run_round.py now ends the options with "--" before the prompt; prompts, schemas, models, tools and the decision rule are unchanged. preregistration-amendment-1.json records the before and after sha256, the smoke results of both clean rooms after the fix, and the gateway rebuild in between. It is recorded before any decision, critic or adjudication of the round. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nstead of selecting The definitive manifest's next version merged (#602) after this round was preregistered against #591; it is the install record on main and has one owner. Before any packet or decision, the round is retargeted: its output is a clean-room audit of that manifest on the 43 slots plus the verified dossiers, with no architecture document of its own. - compare.py: each slot against the manifest row its source_slot names (43 of 43), at #602's merge commit pinned by sha256; verdicts agree, contest, nominates, cross_check, pin_agree, pin_conflict and not_settled, fixed before any decision exists; 21 manifest rows have no slot and are listed as not covered. - run_round.py: the packets state the audit's sampling limits (the first five release assets queried for attestations; one page of check runs, 13 repositories affected) and read a byte-identical copy of the frozen audit observations, so the round no longer depends on the upstream audit's pull request. - freeze.py --check applies the amendment chain (passes; a mutation of compare.py fails it). - Run notes: the 135 empty records left by the usage-limit stop were deleted and are being redone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ation a usage-limit stop ended Usage-limit stops ended the verifier attempts of 13 dossiers after their writers finished; the frozen runner saved them without a verification result and skips saved records. reverify.py re-runs exactly the frozen verification loop (same prompt, model, tools, schema, timeout; at most one repair, none for the two dossiers already repaired once). Recorded before any packet or decision; freeze.py --check passes through the amendment chain. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dment 4 (compare.py defect fix) Both families decided every unit in both orders with critics, and adjudicators ran on the 9 contested slots (GPT-6 Astra at max through OmniRoute; Claude Opus 5.5 at max in safe mode). 177 verified dossiers. selection.json: 35 definitive, 7 measurement, 1 user-pin conflict; contamination audit 0 hits. audit-of-manifest.json against #602 (675bdd5): agree 22, contest 7, nominates 4, cross-check 1, pin agree 1, pin conflict 1, not settled 7; 21 manifest rows not covered. Amendment 4, made with the results known and disclosed as such: compare.py read installs_nothing_extra as a NONE pick, contradicting amendment 2's rule text; the fix turns build-provenance (actions/attest) and dependency-updates (Dependabot) from contest to agree. The pre-fix output is kept beside the fixed one. Three committed dossier copies carry 12-character commit hashes for the secret scanner (README). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…round's records from blind checkouts The decision record states the result against #602 (and unchanged against #620): agree 22, contest 7, nominates 4, memory cross-check, pin agree 1, pin conflict 1, not settled 7. It frames the contests as slot-level picks that did not weigh job overlap with the installed stack, and discloses the unequal live web evidence (the GPT judges' searches through the gateway returned nothing), amendment 4 and the adjudicator-anonymity limit. run-notes.json carries the timeline, the ordering that kept a family's decisions away from the other family's deciders, and usage. blind_checkout.py withholds the round's top-level records and its decision record; the dossiers stay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…read audit An audit of all 2,504 GPT shell commands found no read of the other family's records, but codex listed the user's skills (~/.agents/skills, which an isolated CODEX_HOME does not hide) and its own system skills to the GPT judges. Four slots carry a GPT vote that is not independent of the user's installed tools: browser interaction (agent-browser), the task-scoped skill pack (trailofbits/skills and others), findings publication (codeql) and the runtime-worker slots (codex's own skills). The Claude side ran in safe mode without skills. The web-evidence note now states that the Claude search results are not retained and that the live web names this repository's picks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
08227eb to
5fdeec3
Compare
|
FINDINGS STAND at head I re-read the complete All primary source inputs cited by the seven prior findings retain their exact Git identities. That verdict's five P2 and two P3 findings therefore remain open at this head:
Independent checks bind all 215 PR source paths, 214 unchanged owned Git identities, all 212 registered owned payloads, and every finding's primary input. The complete registry preserves all 9,419 foreign main rows/order, 187 receipts and 26 convergence records. Native source observations retain 18 exit0 and the initial GraphQL alias-conflict exit1, followed by its corrected request; a separate root processor misclassified two commit records as blob records and was corrected before normal verification completion. Whole-source ACCEPT remains held. This is a source/integration read; the main carry does not establish a new blinded trial, runtime role, default, installation or required-CI acceptance. No second Astra invocation or private provider-body read was used for this unchanged finding set. |
…CI of #663 The GPT-6.1 Sol review found four blocking and five non-blocking defects; CodeQL flagged compare.py's URL matching and a repository test found a quoted betterleaks inline-allow marker. Made with the results known and disclosed as such; no outcome status or verdict changes. - assemble.py: an undetermined family's decider picks join the finalists (document ingestion now lists Docling and MinerU, the pair its measurement compares); the run record keeps every dossier call (177 groups). - compare.py parses URLs (output byte-identical); freeze.py rejects an amendment that re-adds a frozen file. - blind_checkout.py withholds every top-level file type of the round folder (assemble.py names the user's pins), with the dossiers kept; the test covers assemble.py and the README. - The decision record corrects the split of the 35 definitive slots (33 agreed, 2 adjudicated), states the dossier revision mix (148 tagged, 23 untagged snapshots, 6 pages) and discloses the runner defect that cost three GPT calls. - The betterleaks dossier copy writes the quoted marker as betterleaks[:]allow. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
5fdeec3 to
7cb9021
Compare
|
Cross-family review and repair (Claude session The review. GPT-6.1 Sol at max, through the OmniRoute gateway, read-only, on a clean checkout of Disposition, all in
Confirmed by the review. Amendment 4's scope holds: reverting it changes only attest and Dependabot. Checks on |
|
Exact-head delta verdict: FINDINGS at Closed: F1's blind-export omission is repaired by the expanded top-level exclusions. F6's split is corrected to 33 family agreements + two four-of-four adjudications, with the pin-conflict agreement counted separately. F7's 148 tagged / 23 untagged / six page sources is supported by the now-explicit rule that every repository of a multi-repository candidate must be tagged. Four medium/P2 findings remain, two with substantial partial repairs:
Custody: all 35 returned native capture exits were 0; a separate pre-command processing syntax failure is preserved. Root verified all 27 complete changed old/new blobs, all 216 PR payload byte/SHA/Git bindings, 202 unchanged OID references and 213 registered bindings. Passing unchanged source evidence and the earlier consequential Astra findings were reused where the affected inputs still match; no new model run or repository/helper test was performed for this delta. Owner-reported checks and historical execution remain separate from this verdict. The post-result amendment 5 is disclosed; it does not turn these corrections into preregistered outcomes or qualify defaults on NativeStack2604. |
|
Closed with a record by the PR triage of 2026-10-07 (the command center's ruling, item review-ns2604-coop-20261007T023012Z (the command center's PR-triage ruling of 2026-10-07; proposal by github-ci-finalize, triage-20261007.json)). Not merged; the branch What it holds: docs/decisions/2026-10-03-definitive-round-audit.md; evidence/artifacts/new-wsl-definitive-round-20261002/ (211 files, 177 candidate dossiers); tools/sota-convergence/blind_checkout.py (modified); tests/test_definitive_round_compare.py (new), tests/test_blind_checkout.py (modified) Superseded by: Overtaken by later manifest decisions: it audits the definitive manifest pinned at 675bdd5 (#602), which 652c15a (#620, the 89-slot layer consensus) had already revised; 3e343ba (#702, layer consensus wave 2) and 1796303 (#713, final architecture round 2 on the repository-quality rule) re-decided it before the NativeStack2604 install. (confidence: low: inference; no landed record cites #663 or adopts its contests or dossiers) Reopen trigger: The definitive manifest is re-audited clean-room by both model families, or a slot contest needs the 177 candidate dossiers (the adoption synthesis and grand-catalog may cite them from this branch at the pinned commit). Reopen with |
Scope
claude -pin safe mode) and GPT-6 Astra at max (codex execwith an isolatedCODEX_HOMEthrough OmniRoute).compare.pysets the outcome against Definitive manifest, next version: the blind GPT round combined with the Claude record (one job per row, 31 definitive, 4 split) #602's manifest.4ced2923(main).claude/definitive-round-20261002before any decision. They are kept reachable asreceipt-revision/860b3480,bd99d15e,e50de81bande9b272d6, and cherry-picked here onto main.lane:foundation.evidence/artifacts/new-wsl-definitive-round-20261002/(new)docs/decisions/2026-10-03-definitive-round-audit.md(new)tests/test_definitive_round_compare.py(new)tools/sota-convergence/blind_checkout.pyandtests/test_blind_checkout.py: the round's records are withheld from blind checkoutsmanifests/evidence.json(registration, in the last commit)Result against the definitive manifest
From
audit-of-manifest.jsonat #602 (675bdd51). The same verdicts hold against main after #620.Limits are in the decision record:
CODEX_HOMEdoes not hide, to the GPT judges. The GPT vote is therefore not independent on 4 slots; the Claude vote is. The slots are browser interaction (agent-browser), the task-scoped skill pack (trailofbits/skills), findings publication (codeql) and the runtime-worker slots (codex). An audit of all 2,504 GPT shell commands found no read of the other family's records.SOTA sources
claude -pwith--safe-mode,--restrictedand--strict-mcp-config;codex execof openai/codex through diegosouzapw/OmniRoute.criteria.txt,decision-rule.txt).Evidence-class table
dossiers/,run-record.jsonselection.json(assemble.py)local_integration(deterministic)python3 .../compare.py→audit-of-manifest.jsonlocal_integrationpython3 .../freeze.py --check→ unchangedlocal_integration(counts from the raw returns)run-notes.json→web_evidenceLocal commands run
Host evidence
Not applicable: no file under
evidence/hosts/changes.Checklist
version comment (no floating tags). (No workflow changes.)
permissions: contents: read(or a narrower, explicitly justified addition). (No workflow changes.)
added without a documented owner.
moved or overwritten).
🤖 Generated with Claude Code