Skip to content

GPT-6 family tiering: record the frozen run's results (S1, gpt-6-sol at medium, for mechanical extraction) - #397

Merged
seathatflowsinourveins merged 4 commits into
mainfrom
claude/gpt6-tiering-results-20260927
Sep 27, 2026
Merged

seathatflowsinourveins merged 4 commits into
mainfrom
claude/gpt6-tiering-results-20260927

Conversation

@seathatflowsinourveins

Copy link
Copy Markdown
Owner

Summary

This records the completed run of the frozen GPT-6 family tiering preregistration
(blueprints/convergence-practice/gpt6-family-tiering-20260926) as a sanitized, re-verifiable receipt, within the
preregistration's own "Evidence classes" and "Privacy boundary" sections.

  • All six arms ran once on 2026-09-27, from 07:09:02Z to 07:28:28Z, on nativestack-5975wx-20260925. They ran
    from a separate worktree at origin/main 50b9579f6722, and every batch completed.
  • analyze.py returned final: true, outcome route_mechanical_extraction: S1, gpt-6-sol at medium. S1
    costs 1,620.0 billed tokens per filing against A0's 1,852.3 (−12.5%). Its micro-F1 is 0.9935 against 0.9942, with
    a paired-bootstrap lower bound of −0.0037 against the −0.02 margin. All six arms were routable.
  • Scope: mechanical, deterministically scored 8-K item extraction only. Generalization to other stages is
    untested, and judgment roles stay on gpt-6-astra at max. This PR routes nothing and changes no stage's
    command.

Changes

  • evidence/artifacts/gpt6-family-tiering-20260927/
    • decision.json: the frozen analyze.py output, copied unchanged. A rerun from this branch's base reproduced
      it byte for byte (SHA-256 555ed71c…).
    • run-record.json: provenance, checkout and pre-run checks, launch and loop exits, and per-arm times. It also
      holds call aggregates from each attempt's call.json (metadata only), the one L1 retry, usage totals, versions,
      two deviations (the slot lock directory and the modes inside each call's Codex home), the coordinator's quota
      reading, the analysis rerun and the privacy check.
    • README.md: the evidence class, the decision and its scope, per-arm quality, tokens and time, the retry,
      the deviations, re-verification, privacy, limits and a records table.
  • blueprints/convergence-practice/gpt6-family-tiering-20260926/README.md: a dated "Result, recorded 2026-09-27"
    section linking the receipt and naming both deviations. No frozen file changed, and run_arm.py verify-frozen
    still passes. experiment.json stays the planned record, which its test pins.
  • manifests/evidence.json: the four files registered through the hot-file protocol, in the last commit. The
    component-matrix and grand-list --write runs produced no change.

Commits:

  • 6bf39d97: the receipt and the result section.
  • 4ad021c9: the fixes for the GPT-6 review's two findings.
  • cfd502c9: the fixes for the independent verifier's findings.
  • 32e9d929: the manifest registration, last. It replaces the earlier registration commit 72c43adc, so the
    manifest carries the fixed files' hashes.

Evidence classes

  • Live provider execution: the six arms. These are hosted GPT-6 models on this host through the native Codex
    sign-in, with Codex's own turn.completed counters. Wall times are shared-host, shared-account observations.
  • Reproduction of the analysis step: analyze.py rerun on the retained private state gave byte-identical
    output. It re-checks the analysis, not the provider calls.
  • Independent observation of the bookkeeping: per-arm usage summed from each attempt's call.json equals the
    decision's totals for all six arms.
  • No upstream test and no synthetic fixture is part of the result.

Findings worth reading

  • The cost gap is mostly cache. Of S1's 232.3 billed tokens per filing below A0, 169.2 is more cached input.
    Without the cache credit, counting input plus output, S1 is still the lowest arm, but only 3.3% below A0. A1, the
    same model and prompts as A0, got no cached input and ranks above A0 on billed tokens. The receipt reports this as
    arithmetic on the decision's aggregates, not as a rule change.
  • The session-record check worked live. One L1 call exited 0 with a valid reply, but its session record held a
    custom_tool_call. That matches the isolation probe's code-mode exec signature, and the --json stream showed
    only an agent_message. The call failed as tool_use, was not scored, and its retry was ok. The frozen command
    is an isolation command, not a tool-free one: it still offers exec, wait and request_user_input, and any
    tool call fails the call.

Deviations

  1. The slot lock directory. README step 2 names the host's shared Codex slot directory, but none existed. The
    landscape sweep's pool defaults to its own <work-dir>/locks (codex_job.py). A dedicated owner-only directory
    holding slot-1 to slot-3 was used instead. It bounded the tiering's own concurrency at 3, with each arm's
    summed slot wait at most 0.002 s. Scoring reads each call's own reply, so it does not depend on the pool. Cached
    input reflects the provider's cache state, which scheduling can influence, and the dedicated pool's effect on it
    was not measured. Without the cache credit, S1 still ranks first (1,861.8 input plus output tokens per filing
    against A0's 1,924.8, a view outside the frozen rule). Wall time could also depend on the pool, and the decision
    needed no tie-break.
  2. Modes inside each call's Codex home. plan.json's state_dir gives the whole state tree, codex-home/
    included, as 0700 directories and 0600 files. The runner's own 314 directories and 629 files match, and it
    creates each codex-home at 0700. It sets no umask and never changes the modes of Codex's own files. Inside
    codex-home, Codex created 5,285 directories at 0775 and 9,513 files at 0664 and 1,329 at 0644. All of them lie
    below 0700 attempt directories, and none is copied into the receipt.

Acceptance

Final state, after both review rounds:

  • python3 blueprints/convergence-practice/gpt6-family-tiering-20260926/run_arm.py verify-frozen: frozen inputs verified
  • python3 scripts/validate.py: passed
  • The coordinator's validation suite (validate, host receipts, catalogs, convergence records, component matrix,
    grand list, ecosystem build check, evidence manifest and others): FAILS=0
  • python3 scripts/build_ecosystem.py --check: passed
  • uv run --no-project --with jsonschema --with pyyaml python -B -m unittest tests.test_gpt6_family_tiering_20260926 -v:
    39 tests OK
  • All 44 unittest modules that reference the touched paths: 2,374 tests OK (23 skipped)

Cross-family review

There was one round, as required. GPT-6 reviewed git diff origin/main...HEAD from the worktree:
cx/gpt-6-astra through the local OmniRoute gateway, effort max, codex-cli 0.157.1, read-only sandbox. It ran
from 08:04:16Z to 08:12:50Z and used 85,093 tokens as Codex reported.

The verdict was DEFECTS FOUND, with two low findings. Both were supported, and both are fixed in 4ad021c9:

  • low | receipt README.md | The README claimed "equals A0's in all 10,000 resamples", but the frozen bootstrap
    records only the point estimate and the 2.5th and 97.5th percentiles. It now states the recorded aggregates: A1
    equals A0 on every quality aggregate, and its point estimate and both bounds are zero.
  • low | receipt README.md and run-record.json | The records claimed "no prompt hash", but decision.json's
    frozen map holds the public prompt.txt template hash. Both now exclude rendered batch prompts and their hashes
    and name the template hash. decision.json stays unchanged.

No finding of the GPT-6 round remains open. No second GPT-6 round was run.

Independent verification

One verifier round reported five low findings, none blocking. All five were supported. The first three are fixed
in cfd502c9; the last two are recorded below as a follow-up and a merge step.

  • low | receipt README.md "Scope" | "the frozen, tool-free codex exec command" misstated the condition. It now
    says the frozen isolation command still offers exec, wait and request_user_input, that any tool call fails
    the call and its reply is not scored, and links the preregistration's "Why the overrides".
  • low | receipt README.md, run-record.json deviations[0], this body | "Token counts do not depend on the pool;
    only wall time could" overclaimed. All three now say cached input reflects provider cache state, which scheduling
    can influence, that the pool's effect was not measured, and that S1 still ranks first without the cache credit.
  • low | run-record.json, receipt README.md, preregistration Result section | The codex-home modes were a
    second, undeclared deviation from plan.json's state_dir. run-record.json gains deviations[1], and its
    privacy reading, the receipt's run record, privacy section and records table, and the Result section ("Two
    deviations") now match.
  • low | docs/decisions/2026-09-26-codex-worker-lane.md:80-81 | That record says A0 "has not run", which this PR
    makes stale. It is left unedited here and listed under Follow-ups.
  • low | the earlier handoff's rebase note | It named the wrong origin/main tip. The corrected note is under
    "Before merge".

Before merge

This branch is not rebased. Its base is 50b9579f6722. The local origin/main is 5f3a7c21, two commits past
it: b9abcc5f (#387) and 5f3a7c21 (#388). Both change manifests/evidence.json. #387 also changes
tools/sota-convergence/landscape-sweep/codex_job.py. At 5f3a7c21 it still defaults its slot pool to
<work-dir>/locks (docstring line 32, settings() lines 150-151), so the deviation text holds after a rebase. No
frozen file of the preregistration changed on main. At merge, follow the hot-file protocol (docs/lanes.md): take
main's manifests/evidence.json, re-register the four files, run scripts/component_matrix.py --write and
scripts/new_host_grand_list.py --write, then scripts/validate.py.

Follow-ups

  • docs/decisions/2026-09-26-codex-worker-lane.md:80-81 calls A0 a control arm "which has not run". After merge, a
    later change adds a dated one-line update there, in the style of
    docs/decisions/2026-09-25-workstation-sota-refresh.md ("Update (2026-09-25, after the cutover)."), pointing
    to evidence/artifacts/gpt6-family-tiering-20260927/. Its lines 201-202 and
    adoption/templates/codex.stack-worker.config.toml:13-14 stay accurate, because the tiering changes stage
    commands, not the profile.
  • A future preregistration that reuses this runner can start the Codex child with an owner-only umask, so Codex's
    own files match a 0700/0600 plan. Python's subprocess.Popen takes a umask argument for exactly this (Python
    3.9 and later). The frozen runner stays unchanged, because its SHA-256 is in plan.json's frozen_inputs.

SOTA sources

  • The frozen preregistration, the source of the protocol, rule and code:
    blueprints/convergence-practice/gpt6-family-tiering-20260926/plan.json (SHA-256 455e38df…), with run_arm.py
    and analyze.py at the SHA-256 values in its frozen_inputs. The results are that code's own output.
  • The same preregistration for this round's fixes: README.md "Why the overrides" (the tools that remain with the
    overrides, and that any tool call fails the call), isolation-probe/results.json (cases isolated-code-mode and
    isolated-plain-luna), plan.json state_dir (0700/0600 for the whole tree, codex-home/ included), and
    run_arm.py (codex_home.mkdir(mode=0o700), no umask call, and its only chmod setting its own files to 0600).
  • li26's reused scorers, unchanged: blueprints/convergence-practice/local-inference-latest-20260926/analyze.py
    and eval_arm.py (hashes in the same frozen_inputs).
  • openai/codex rust-v0.157.1, commit 36650394c5b38c2990ccf2a3457165ca3e9d9726: the CLI that ran every call, and
    the source of its exec --json event schema and turn.completed usage counters. The receipt is
    evidence/receipts/codex-01571-qualification-20260926.json.
  • docs/token-practice.md: Codex's total is input plus output, and cached input and reasoning are subsets.
  • docs/acceptance-evidence-policy.md and docs/lanes.md (hot-file protocol), for evidence classes and
    registration.
  • tools/sota-convergence/landscape-sweep/codex_job.py at 50b9579f (docstring and settings()), and again at
    5f3a7c21 (docstring line 32, settings() lines 150-151): the default <work-dir>/locks slot pool behind the
    first deviation.
  • Python subprocess.Popen, umask parameter (https://docs.python.org/3/library/subprocess.html#subprocess.Popen,
    added in 3.9): the mechanism named in the umask follow-up. Nothing here uses it.

Lane

Recommended label: lane:foundation.

🤖 Generated with Claude Code

Scout and others added 4 commits September 27, 2026 07:41
…receipt

The frozen preregistration blueprints/convergence-practice/gpt6-family-tiering-20260926
ran once per arm on 2026-09-27 (07:09:02Z to 07:28:28Z) on
nativestack-5975wx-20260925, from a separate worktree at 50b9579.
analyze.py returned final true, outcome route_mechanical_extraction:
S1, gpt-6-sol at medium, 1,620.0 billed tokens per filing against A0's
1,852.3, micro-F1 0.9935 against 0.9942 (bootstrap lower bound -0.0037).

evidence/artifacts/gpt6-family-tiering-20260927/:
- decision.json: the analyze.py output, copied unchanged. A rerun from
  this checkout reproduced it byte for byte (SHA-256 555ed71c...).
- run-record.json: provenance, checkout and pre-run checks, launch and
  loop exits, per-arm times from each arm's runs.jsonl, call aggregates
  from each attempt's call.json (metadata only), the one L1 retry
  (tool_use caught by the session-record check), usage totals, versions,
  the slot-lock-directory deviation, the coordinator's quota reading and
  the privacy check.
- README.md: evidence class (live provider execution), decision and its
  scope, per-arm quality, tokens and time, the retry, re-verification,
  privacy and limits, and a records table.

The preregistration README gains a dated result section linking the
receipt. No frozen file changes: run_arm.py verify-frozen still passes.
experiment.json stays the planned record that its test pins.

Scope: mechanical, deterministically scored 8-K item extraction only;
generalization to other stages is untested; judgment roles stay on
gpt-6-astra at max; no stage's command changes here.

Sources: the preregistration's plan.json, analyze.py and run_arm.py
(frozen hashes in plan.json), li26's scorers, openai/codex rust-v0.157.1
(36650394c5b3), docs/token-practice.md, docs/acceptance-evidence-policy.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
One cross-family review round (gpt-6-astra at max through codex-cli
0.157.1, read-only) returned DEFECTS FOUND with two low findings, both
supported by the files:

- README said A1's micro-F1 equals A0's in all 10,000 resamples. The
  frozen paired_bootstrap records only the point estimate and the 2.5th
  and 97.5th percentiles, so zero bounds do not show every resample.
  Now: A1 equals A0 on every quality aggregate in the decision, and its
  point estimate and both 95% bounds are exactly zero.
- README and run-record.json said neither record contains a prompt hash,
  but decision.json's frozen map holds the public prompt.txt template's
  SHA-256. Both now exclude rendered batch prompts and their hashes and
  name the template hash. decision.json stays unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
One independent verification round reported five low findings. Three
change files here; the other two are PR-body follow-ups.

- Scope: the Covered bullet called the frozen command tool-free. With
  the overrides it still offers exec, wait and request_user_input, and
  any tool call fails the call (preregistration README, "Why the
  overrides"; isolation-probe/results.json; the L1 b016 retry). The
  bullet now says so and links that section.
- Slot-lock deviation: the effect claimed token counts cannot depend on
  the pool. Cached input reflects provider cache state, which
  scheduling can influence, and the pool's effect on it was not
  measured. The README and run-record.json now say that, keep the
  scoring sentence, and point to the no-cache-credit view (S1 first at
  1,861.8 input plus output per filing against A0's 1,924.8, outside
  the frozen rule).
- A second deviation: plan.json state_dir gives the whole state tree,
  codex-home included, as 0700/0600. Inside each 0700 codex-home, Codex
  created directories at 0775 and files at 0664 and 0644 (the recorded
  privacy_check counts). run_arm.py sets no umask and never re-modes
  Codex's files. run-record.json gains deviations[1], its privacy
  reading and the receipt's privacy section and records table point to
  it, and the preregistration's Result section now says two deviations.

decision.json is unchanged (SHA-256 555ed71c...). No frozen file changed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Hot-file protocol (docs/lanes.md): host_receipts.register_file for the
three new files under evidence/artifacts/gpt6-family-tiering-20260927/
and the changed preregistration README, then component_matrix.py
--write and new_host_grand_list.py --write, which changed nothing else.
This is the branch's last commit and its only manifests/evidence.json
edit; it replaces the earlier registration commit (72c43adc) so that
the manifest carries the hashes of the verifier fixes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant