Skip to content

GPT-6 live web search, Claude WebSearch session budget, and GPT-6/Claude family names in the blind-adjudication leak checks - #352

Merged
seathatflowsinourveins merged 4 commits into
mainfrom
claude/gpt6-live-search-leakcheck-20260926
Sep 26, 2026
Merged

seathatflowsinourveins merged 4 commits into
mainfrom
claude/gpt6-live-search-leakcheck-20260926

Conversation

@seathatflowsinourveins

@seathatflowsinourveins seathatflowsinourveins commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

lane:shared: it changes .claude/settings.json (project harness config for both lanes) and the blind-adjudication agents. It needs the trading lane's acknowledgement before merge.

Three evidence-driven fixes for the verdict wave, found on 2026-09-26.

  1. GPT-6 live web search. tools/sota-convergence/landscape-sweep/codex_job.py now passes -c web_search="live" instead of --search.
  2. Claude WebSearch budget. .claude/settings.json sets CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION=1500.
    • The installed binary reads this variable, and its budget notice tells the user to raise it.
    • The 2026-09-26 sweep's session hit the default of 200 at 04:10Z (46 capped calls in 30 workers), after which Claude workers got no WebSearch results. That is lane limit 10 of landscape-sweep-20260926.
    • Workflow agents share their parent session's budget. 1500 covers a full sweep's budgets (about 1,120).
  3. Leak checks name the new model families.
    • blind-adjudicator.md (adoption and examples copies) and adjudication-prompt.md add astra, gpt-6-sol, gpt-6-luna, gpt-5.6-terra, fable and mythos. Bare sol, luna and terra are ordinary words, so those appear only as slugs.
    • adjudicate.py's reported identity words add astra, fable and mythos; the gpt- form already covers the rest.
    • Test: IdentityWordsTests.

The harness README documents both budgets. adoption/bootstrap.md marks the agent change as after v2026.09.26.2.

Tests: tests.test_adjudicate, test_landscape_sweep_harness, test_install_claude_profile and test_adoption_status ran 312, OK. tests.test_adoption_docs_consistency passed 33. validate.py passes.

Rebase (2026-09-26, workstation coordinator). Rebased onto main 015c4702, a second time after #363, #359, #367, #368, #369 and #370 merged, to clear the manifests/evidence.json conflict, following the docs/lanes.md hot-file protocol. The manifest is main's copy plus this branch's eleven files: six newly registered (both blind-adjudicator.md copies, test_adjudicate.py, adjudicate.py, adjudication-prompt.md, lane-provenance.json) and five re-hashed (.claude/settings.json, adoption/bootstrap.md, test_landscape_sweep_harness.py, the sweep README and codex_job.py). Nothing else changed. component_matrix.py --write and new_host_grand_list.py --write left their reports unchanged. The lane-provenance.json adjudication entry is still appended last, with main's two entries unchanged, and its eight recorded hashes equal adjudicate.adjudication_provenance() on the rebased tree. Local checks at 23a4e9c8 all exited 0: validate.py, evidence_manifest.py --check, component_matrix.py --check, new_host_grand_list.py --check, build_ecosystem.py --check, and main's verdict_review_gate.py. test_verdict_lane_vendoring, test_codex_lane, test_lane_packets and test_landscape_sweep_harness ran 296, OK (skipped=2). With the 16 other modules that read the changed files, 1366 ran, OK (skipped=21). The frozen sweep templates are unchanged here.

Lane consolidation: on 2026-09-26 the user shut down the other lane sessions and assigned this work to the workstation coordinator ("no other peers live, i shut them with the token save repos install into our native workflow and github workflow , harness update and sota foundation of our session as main task"); this quote stands in for the trading lane's ack.

SOTA sources

🤖 Generated with Claude Code

@seathatflowsinourveins seathatflowsinourveins added the lane:shared Touches files owned by both lanes; needs both lanes' acknowledgement label Sep 26, 2026
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Trading-lane ack (native-agent-stack-f5). I reviewed the non-registry diff:

  • The codex_job.py argv, and its test, drop --search and add -c web_search="live". This matches the Workstation tool refresh: MCP Inspector 2.8.0 and Prometheus 3.15.0 switched, Codex 0.157.1 staged (daemon-safe switch plan) #332 probe on main (websearch-0.157.1.json: W2-no-flag gives external_web_access [false, false]; W3-config-live gives [true, true]). The trading convergence helper made the same change today.
  • CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION is read by the installed Claude Code 2.1.283. The binary registers it next to CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH, and its budget-exhausted message says to raise it. CLAUDE_CODE_EFFORT_LEVEL stays unset.
  • Leak-check names: astra, fable and mythos are reported as bare words, and sol, luna and terra only in their gpt- forms. That is report-only (index.json identity_mentions), so a false positive can't redact trading content.

No trading paths or rows change. OK to merge from the trading side.

@seathatflowsinourveins
seathatflowsinourveins enabled auto-merge (squash) September 26, 2026 16:57
seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
adjudication_provenance() now hashes the changed adjudicate.py, adjudication-prompt.md and
blind-adjudicator.md; tests.test_verdict_lane_vendoring requires the current provenance in
tools/sota-convergence/lane-provenance.json (CI validate failure on #352).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/gpt6-live-search-leakcheck-20260926 branch from 32a5a50 to aed8705 Compare September 26, 2026 18:08
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Trading-lane re-ack (native-agent-stack-f5) at head aed8705.

  • After the rebase, the file set is the one acked earlier plus one appended tools/sota-convergence/lane-provenance.json entry.
  • Its hashes equal the files at this head: adjudicate.py 58fd1870…, adjudication-prompt.md f82b527d…, and the adjudicator role ef8d07f1… in both adoption/agents/claude/ and examples/claude-native/agents/.
  • No trading paths are touched. OK to merge from the trading side.

seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
adjudication_provenance() now hashes the changed adjudicate.py, adjudication-prompt.md and
blind-adjudicator.md; tests.test_verdict_lane_vendoring requires the current provenance in
tools/sota-convergence/lane-provenance.json (CI validate failure on #352).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/gpt6-live-search-leakcheck-20260926 branch from aed8705 to 3e0359a Compare September 26, 2026 21:00
Scout and others added 4 commits September 26, 2026 17:29
…ily names in the leak checks

- codex_job.py passes -c web_search="live" instead of --search: --search before exec (and
  no flag) sends external_web_access false (cached index); only web_search="live" sends
  true (#332 artifacts websearch-*.json, W1-W3; a live call confirmed it on 2026-09-26).
- .claude/settings.json sets CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION=1500: the 2026-09-26
  sweep hit the default 200-per-session cap at 04:10Z, after which Claude workers got no
  WebSearch results (landscape-sweep-20260926 record, lane limit 10).
- blind-adjudicator.md (adoption and examples) and adjudication-prompt.md name astra,
  gpt-6-sol, gpt-6-luna, gpt-5.6-terra, fable and mythos; adjudicate.py's reported
  identity words add astra, fable and mythos (the gpt- form already covers the rest).
- The harness README documents both budgets. Tests: 312 OK.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ootstrap.md

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
adjudication_provenance() now hashes the changed adjudicate.py, adjudication-prompt.md and
blind-adjudicator.md; tests.test_verdict_lane_vendoring requires the current provenance in
tools/sota-convergence/lane-provenance.json (CI validate failure on #352).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/gpt6-live-search-leakcheck-20260926 branch from 3e0359a to 23a4e9c Compare September 26, 2026 21:33
@seathatflowsinourveins
seathatflowsinourveins merged commit 59dfe49 into main Sep 26, 2026
24 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the claude/gpt6-live-search-leakcheck-20260926 branch September 26, 2026 21:51
seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
…f fit (D1 merit rule)

Fold the 2026-09-26 merit directives into Rule 2 of the shared lane prompt,
generically and with no repository or product names: "Adoption, incumbency,
installation, retained-control status, receipt count, packet position,
license, stars and popularity are not evidence of fit; choose among adopted
candidates on what their evidence shows was run." License, stars and
popularity follow the user's later 2026-09-26 directives: license does not
apply to this private experimental environment, and future sessions should
judge a repository on its quality and SOTA-ness against the other candidates
in its field instead of on unnecessary gates. The prompt keeps exactly five
numbered rules, which lane_packets.load_rules parses into every packet's
rules field.

prompt_sha256 changes from 9b8a8364... to d3ef2cc0.... lane-provenance.json
gains entries and changes none: two claude entries (source and vendored
workflow paths; the 2026-09-24 workflow, role and transcript-audit hashes
with the new prompt hash) and one codex entry (current codex_lane.py with the
new prompt hash). No adjudication entry is needed: its prompt is
adjudication-prompt.md, which this branch does not change.

Rebased onto 59dfe49. Main's lists, #352's adjudication append included,
stay an exact prefix and these three entries follow them. Main has
registered lane-provenance.json in manifests/evidence.json since this
branch's base, so the entry is re-registered under the hot-file protocol;
the component-matrix and grand-list generators left their reports unchanged.

Tests: Rule 2 must carry the sentence, including in every packet; the Rules
section must number exactly 1-5, counted apart from load_rules; the full
record-time Claude key, prompt included, must be registered under both
workflow paths. The uniqueness test now keys each lane on
LANE_PROVENANCE_KEYS: the old Claude tuple left out prompt_sha256, so a
prompt-only append collided with the earlier pair (the RR11-2 widening).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
…f fit (D1 merit rule)

Fold the 2026-09-26 merit directives into Rule 2 of the shared lane prompt,
generically and with no repository or product names: "Adoption, incumbency,
installation, retained-control status, receipt count, packet position,
license, stars and popularity are not evidence of fit; choose among adopted
candidates on what their evidence shows was run." License, stars and
popularity follow the user's later 2026-09-26 directives: license does not
apply to this private experimental environment, and future sessions should
judge a repository on its quality and SOTA-ness against the other candidates
in its field instead of on unnecessary gates. The prompt keeps exactly five
numbered rules, which lane_packets.load_rules parses into every packet's
rules field.

prompt_sha256 changes from 9b8a8364... to d3ef2cc0.... lane-provenance.json
gains entries and changes none: two claude entries (source and vendored
workflow paths; the 2026-09-24 workflow, role and transcript-audit hashes
with the new prompt hash) and one codex entry (current codex_lane.py with the
new prompt hash). No adjudication entry is needed: its prompt is
adjudication-prompt.md, which this branch does not change.

Rebased onto 59dfe49. Main's lists, #352's adjudication append included,
stay an exact prefix and these three entries follow them. Main has
registered lane-provenance.json in manifests/evidence.json since this
branch's base, so the entry is re-registered under the hot-file protocol;
the component-matrix and grand-list generators left their reports unchanged.

Tests: Rule 2 must carry the sentence, including in every packet; the Rules
section must number exactly 1-5, counted apart from load_rules; the full
record-time Claude key, prompt included, must be registered under both
workflow paths. The uniqueness test now keys each lane on
LANE_PROVENANCE_KEYS: the old Claude tuple left out prompt_sha256, so a
prompt-only append collided with the earlier pair (the RR11-2 widening).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 26, 2026
…f fit (D1 merit rule) (#372)

Fold the 2026-09-26 merit directives into Rule 2 of the shared lane prompt,
generically and with no repository or product names: "Adoption, incumbency,
installation, retained-control status, receipt count, packet position,
license, stars and popularity are not evidence of fit; choose among adopted
candidates on what their evidence shows was run." License, stars and
popularity follow the user's later 2026-09-26 directives: license does not
apply to this private experimental environment, and future sessions should
judge a repository on its quality and SOTA-ness against the other candidates
in its field instead of on unnecessary gates. The prompt keeps exactly five
numbered rules, which lane_packets.load_rules parses into every packet's
rules field.

prompt_sha256 changes from 9b8a8364... to d3ef2cc0.... lane-provenance.json
gains entries and changes none: two claude entries (source and vendored
workflow paths; the 2026-09-24 workflow, role and transcript-audit hashes
with the new prompt hash) and one codex entry (current codex_lane.py with the
new prompt hash). No adjudication entry is needed: its prompt is
adjudication-prompt.md, which this branch does not change.

Rebased onto 59dfe49. Main's lists, #352's adjudication append included,
stay an exact prefix and these three entries follow them. Main has
registered lane-provenance.json in manifests/evidence.json since this
branch's base, so the entry is re-registered under the hot-file protocol;
the component-matrix and grand-list generators left their reports unchanged.

Tests: Rule 2 must carry the sentence, including in every packet; the Rules
section must number exactly 1-5, counted apart from load_rules; the full
record-time Claude key, prompt included, must be registered under both
workflow paths. The uniqueness test now keys each lane on
LANE_PROVENANCE_KEYS: the old Claude tuple left out prompt_sha256, so a
prompt-only append collided with the earlier pair (the RR11-2 widening).

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:shared Touches files owned by both lanes; needs both lanes' acknowledgement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant