Repository navigation
GPT-6 family tiering: record the frozen run's results (S1, gpt-6-sol at medium, for mechanical extraction) - #397
Merged
seathatflowsinourveins merged 4 commits intoSep 27, 2026
Conversation
…receipt The frozen preregistration blueprints/convergence-practice/gpt6-family-tiering-20260926 ran once per arm on 2026-09-27 (07:09:02Z to 07:28:28Z) on nativestack-5975wx-20260925, from a separate worktree at 50b9579. analyze.py returned final true, outcome route_mechanical_extraction: S1, gpt-6-sol at medium, 1,620.0 billed tokens per filing against A0's 1,852.3, micro-F1 0.9935 against 0.9942 (bootstrap lower bound -0.0037). evidence/artifacts/gpt6-family-tiering-20260927/: - decision.json: the analyze.py output, copied unchanged. A rerun from this checkout reproduced it byte for byte (SHA-256 555ed71c...). - run-record.json: provenance, checkout and pre-run checks, launch and loop exits, per-arm times from each arm's runs.jsonl, call aggregates from each attempt's call.json (metadata only), the one L1 retry (tool_use caught by the session-record check), usage totals, versions, the slot-lock-directory deviation, the coordinator's quota reading and the privacy check. - README.md: evidence class (live provider execution), decision and its scope, per-arm quality, tokens and time, the retry, re-verification, privacy and limits, and a records table. The preregistration README gains a dated result section linking the receipt. No frozen file changes: run_arm.py verify-frozen still passes. experiment.json stays the planned record that its test pins. Scope: mechanical, deterministically scored 8-K item extraction only; generalization to other stages is untested; judgment roles stay on gpt-6-astra at max; no stage's command changes here. Sources: the preregistration's plan.json, analyze.py and run_arm.py (frozen hashes in plan.json), li26's scorers, openai/codex rust-v0.157.1 (36650394c5b3), docs/token-practice.md, docs/acceptance-evidence-policy.md. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
One cross-family review round (gpt-6-astra at max through codex-cli 0.157.1, read-only) returned DEFECTS FOUND with two low findings, both supported by the files: - README said A1's micro-F1 equals A0's in all 10,000 resamples. The frozen paired_bootstrap records only the point estimate and the 2.5th and 97.5th percentiles, so zero bounds do not show every resample. Now: A1 equals A0 on every quality aggregate in the decision, and its point estimate and both 95% bounds are exactly zero. - README and run-record.json said neither record contains a prompt hash, but decision.json's frozen map holds the public prompt.txt template's SHA-256. Both now exclude rendered batch prompts and their hashes and name the template hash. decision.json stays unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
One independent verification round reported five low findings. Three change files here; the other two are PR-body follow-ups. - Scope: the Covered bullet called the frozen command tool-free. With the overrides it still offers exec, wait and request_user_input, and any tool call fails the call (preregistration README, "Why the overrides"; isolation-probe/results.json; the L1 b016 retry). The bullet now says so and links that section. - Slot-lock deviation: the effect claimed token counts cannot depend on the pool. Cached input reflects provider cache state, which scheduling can influence, and the pool's effect on it was not measured. The README and run-record.json now say that, keep the scoring sentence, and point to the no-cache-credit view (S1 first at 1,861.8 input plus output per filing against A0's 1,924.8, outside the frozen rule). - A second deviation: plan.json state_dir gives the whole state tree, codex-home included, as 0700/0600. Inside each 0700 codex-home, Codex created directories at 0775 and files at 0664 and 0644 (the recorded privacy_check counts). run_arm.py sets no umask and never re-modes Codex's files. run-record.json gains deviations[1], its privacy reading and the receipt's privacy section and records table point to it, and the preregistration's Result section now says two deviations. decision.json is unchanged (SHA-256 555ed71c...). No frozen file changed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Hot-file protocol (docs/lanes.md): host_receipts.register_file for the three new files under evidence/artifacts/gpt6-family-tiering-20260927/ and the changed preregistration README, then component_matrix.py --write and new_host_grand_list.py --write, which changed nothing else. This is the branch's last commit and its only manifests/evidence.json edit; it replaces the earlier registration commit (72c43adc) so that the manifest carries the hashes of the verifier fixes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
deleted the
claude/gpt6-tiering-results-20260927
branch
September 27, 2026 12:03
This was referenced Sep 27, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This records the completed run of the frozen GPT-6 family tiering preregistration
(
blueprints/convergence-practice/gpt6-family-tiering-20260926) as a sanitized, re-verifiable receipt, within thepreregistration's own "Evidence classes" and "Privacy boundary" sections.
nativestack-5975wx-20260925. They ranfrom a separate worktree at
origin/main50b9579f6722, and every batch completed.analyze.pyreturnedfinal: true, outcomeroute_mechanical_extraction: S1,gpt-6-solatmedium. S1costs 1,620.0 billed tokens per filing against A0's 1,852.3 (−12.5%). Its micro-F1 is 0.9935 against 0.9942, with
a paired-bootstrap lower bound of −0.0037 against the −0.02 margin. All six arms were routable.
untested, and judgment roles stay on
gpt-6-astraatmax. This PR routes nothing and changes no stage'scommand.
Changes
evidence/artifacts/gpt6-family-tiering-20260927/decision.json: the frozenanalyze.pyoutput, copied unchanged. A rerun from this branch's base reproducedit byte for byte (SHA-256
555ed71c…).run-record.json: provenance, checkout and pre-run checks, launch and loop exits, and per-arm times. It alsoholds call aggregates from each attempt's
call.json(metadata only), the one L1 retry, usage totals, versions,two deviations (the slot lock directory and the modes inside each call's Codex home), the coordinator's quota
reading, the analysis rerun and the privacy check.
README.md: the evidence class, the decision and its scope, per-arm quality, tokens and time, the retry,the deviations, re-verification, privacy, limits and a records table.
blueprints/convergence-practice/gpt6-family-tiering-20260926/README.md: a dated "Result, recorded 2026-09-27"section linking the receipt and naming both deviations. No frozen file changed, and
run_arm.py verify-frozenstill passes.
experiment.jsonstays the planned record, which its test pins.manifests/evidence.json: the four files registered through the hot-file protocol, in the last commit. Thecomponent-matrix and grand-list
--writeruns produced no change.Commits:
6bf39d97: the receipt and the result section.4ad021c9: the fixes for the GPT-6 review's two findings.cfd502c9: the fixes for the independent verifier's findings.32e9d929: the manifest registration, last. It replaces the earlier registration commit72c43adc, so themanifest carries the fixed files' hashes.
Evidence classes
sign-in, with Codex's own
turn.completedcounters. Wall times are shared-host, shared-account observations.analyze.pyrerun on the retained private state gave byte-identicaloutput. It re-checks the analysis, not the provider calls.
call.jsonequals thedecision's totals for all six arms.
Findings worth reading
Without the cache credit, counting input plus output, S1 is still the lowest arm, but only 3.3% below A0. A1, the
same model and prompts as A0, got no cached input and ranks above A0 on billed tokens. The receipt reports this as
arithmetic on the decision's aggregates, not as a rule change.
custom_tool_call. That matches the isolation probe's code-modeexecsignature, and the--jsonstream showedonly an
agent_message. The call failed astool_use, was not scored, and its retry wasok. The frozen commandis an isolation command, not a tool-free one: it still offers
exec,waitandrequest_user_input, and anytool call fails the call.
Deviations
landscape sweep's pool defaults to its own
<work-dir>/locks(codex_job.py). A dedicated owner-only directoryholding
slot-1toslot-3was used instead. It bounded the tiering's own concurrency at 3, with each arm'ssummed slot wait at most 0.002 s. Scoring reads each call's own reply, so it does not depend on the pool. Cached
input reflects the provider's cache state, which scheduling can influence, and the dedicated pool's effect on it
was not measured. Without the cache credit, S1 still ranks first (1,861.8 input plus output tokens per filing
against A0's 1,924.8, a view outside the frozen rule). Wall time could also depend on the pool, and the decision
needed no tie-break.
plan.json'sstate_dirgives the whole state tree,codex-home/included, as 0700 directories and 0600 files. The runner's own 314 directories and 629 files match, and it
creates each
codex-homeat 0700. It sets no umask and never changes the modes of Codex's own files. Insidecodex-home, Codex created 5,285 directories at 0775 and 9,513 files at 0664 and 1,329 at 0644. All of them liebelow 0700 attempt directories, and none is copied into the receipt.
Acceptance
Final state, after both review rounds:
python3 blueprints/convergence-practice/gpt6-family-tiering-20260926/run_arm.py verify-frozen:frozen inputs verifiedpython3 scripts/validate.py: passedgrand list, ecosystem build check, evidence manifest and others):
FAILS=0python3 scripts/build_ecosystem.py --check: passeduv run --no-project --with jsonschema --with pyyaml python -B -m unittest tests.test_gpt6_family_tiering_20260926 -v:39 tests OK
Cross-family review
There was one round, as required. GPT-6 reviewed
git diff origin/main...HEADfrom the worktree:cx/gpt-6-astrathrough the local OmniRoute gateway, effortmax,codex-cli 0.157.1, read-only sandbox. It ranfrom 08:04:16Z to 08:12:50Z and used 85,093 tokens as Codex reported.
The verdict was DEFECTS FOUND, with two low findings. Both were supported, and both are fixed in
4ad021c9:README.md| The README claimed "equals A0's in all 10,000 resamples", but the frozen bootstraprecords only the point estimate and the 2.5th and 97.5th percentiles. It now states the recorded aggregates: A1
equals A0 on every quality aggregate, and its point estimate and both bounds are zero.
README.mdandrun-record.json| The records claimed "no prompt hash", butdecision.json'sfrozen map holds the public
prompt.txttemplate hash. Both now exclude rendered batch prompts and their hashesand name the template hash.
decision.jsonstays unchanged.No finding of the GPT-6 round remains open. No second GPT-6 round was run.
Independent verification
One verifier round reported five low findings, none blocking. All five were supported. The first three are fixed
in
cfd502c9; the last two are recorded below as a follow-up and a merge step.README.md"Scope" | "the frozen, tool-freecodex execcommand" misstated the condition. It nowsays the frozen isolation command still offers
exec,waitandrequest_user_input, that any tool call failsthe call and its reply is not scored, and links the preregistration's "Why the overrides".
README.md,run-record.jsondeviations[0], this body | "Token counts do not depend on the pool;only wall time could" overclaimed. All three now say cached input reflects provider cache state, which scheduling
can influence, that the pool's effect was not measured, and that S1 still ranks first without the cache credit.
run-record.json, receiptREADME.md, preregistration Result section | Thecodex-homemodes were asecond, undeclared deviation from
plan.json'sstate_dir.run-record.jsongainsdeviations[1], and itsprivacy reading, the receipt's run record, privacy section and records table, and the Result section ("Two
deviations") now match.
docs/decisions/2026-09-26-codex-worker-lane.md:80-81| That record says A0 "has not run", which this PRmakes stale. It is left unedited here and listed under Follow-ups.
origin/maintip. The corrected note is under"Before merge".
Before merge
This branch is not rebased. Its base is
50b9579f6722. The localorigin/mainis5f3a7c21, two commits pastit:
b9abcc5f(#387) and5f3a7c21(#388). Both changemanifests/evidence.json. #387 also changestools/sota-convergence/landscape-sweep/codex_job.py. At5f3a7c21it still defaults its slot pool to<work-dir>/locks(docstring line 32,settings()lines 150-151), so the deviation text holds after a rebase. Nofrozen file of the preregistration changed on
main. At merge, follow the hot-file protocol (docs/lanes.md): takemain'smanifests/evidence.json, re-register the four files, runscripts/component_matrix.py --writeandscripts/new_host_grand_list.py --write, thenscripts/validate.py.Follow-ups
docs/decisions/2026-09-26-codex-worker-lane.md:80-81calls A0 a control arm "which has not run". After merge, alater change adds a dated one-line update there, in the style of
docs/decisions/2026-09-25-workstation-sota-refresh.md("Update (2026-09-25, after the cutover)."), pointingto
evidence/artifacts/gpt6-family-tiering-20260927/. Its lines 201-202 andadoption/templates/codex.stack-worker.config.toml:13-14stay accurate, because the tiering changes stagecommands, not the profile.
own files match a 0700/0600 plan. Python's
subprocess.Popentakes aumaskargument for exactly this (Python3.9 and later). The frozen runner stays unchanged, because its SHA-256 is in
plan.json'sfrozen_inputs.SOTA sources
blueprints/convergence-practice/gpt6-family-tiering-20260926/plan.json(SHA-256455e38df…), withrun_arm.pyand
analyze.pyat the SHA-256 values in itsfrozen_inputs. The results are that code's own output.README.md"Why the overrides" (the tools that remain with theoverrides, and that any tool call fails the call),
isolation-probe/results.json(casesisolated-code-modeandisolated-plain-luna),plan.jsonstate_dir(0700/0600 for the whole tree,codex-home/included), andrun_arm.py(codex_home.mkdir(mode=0o700), no umask call, and its only chmod setting its own files to 0600).blueprints/convergence-practice/local-inference-latest-20260926/analyze.pyand
eval_arm.py(hashes in the samefrozen_inputs).rust-v0.157.1, commit36650394c5b38c2990ccf2a3457165ca3e9d9726: the CLI that ran every call, andthe source of its
exec --jsonevent schema andturn.completedusage counters. The receipt isevidence/receipts/codex-01571-qualification-20260926.json.docs/token-practice.md: Codex's total is input plus output, and cached input and reasoning are subsets.docs/acceptance-evidence-policy.mdanddocs/lanes.md(hot-file protocol), for evidence classes andregistration.
tools/sota-convergence/landscape-sweep/codex_job.pyat50b9579f(docstring andsettings()), and again at5f3a7c21(docstring line 32,settings()lines 150-151): the default<work-dir>/locksslot pool behind thefirst deviation.
subprocess.Popen,umaskparameter (https://docs.python.org/3/library/subprocess.html#subprocess.Popen,added in 3.9): the mechanism named in the umask follow-up. Nothing here uses it.
Lane
Recommended label:
lane:foundation.🤖 Generated with Claude Code