Repository navigation
GPT-6 live web search, Claude WebSearch session budget, and GPT-6/Claude family names in the blind-adjudication leak checks - #352
Merged
seathatflowsinourveins merged 4 commits intoSep 26, 2026
Conversation
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Owner
Author
|
Trading-lane ack (native-agent-stack-f5). I reviewed the non-registry diff:
No trading paths or rows change. OK to merge from the trading side. |
seathatflowsinourveins
enabled auto-merge (squash)
September 26, 2026 16:57
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
adjudication_provenance() now hashes the changed adjudicate.py, adjudication-prompt.md and blind-adjudicator.md; tests.test_verdict_lane_vendoring requires the current provenance in tools/sota-convergence/lane-provenance.json (CI validate failure on #352). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
force-pushed
the
claude/gpt6-live-search-leakcheck-20260926
branch
from
September 26, 2026 18:08
32a5a50 to
aed8705
Compare
Owner
Author
|
Trading-lane re-ack (native-agent-stack-f5) at head aed8705.
|
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
adjudication_provenance() now hashes the changed adjudicate.py, adjudication-prompt.md and blind-adjudicator.md; tests.test_verdict_lane_vendoring requires the current provenance in tools/sota-convergence/lane-provenance.json (CI validate failure on #352). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
force-pushed
the
claude/gpt6-live-search-leakcheck-20260926
branch
from
September 26, 2026 21:00
aed8705 to
3e0359a
Compare
…ily names in the leak checks - codex_job.py passes -c web_search="live" instead of --search: --search before exec (and no flag) sends external_web_access false (cached index); only web_search="live" sends true (#332 artifacts websearch-*.json, W1-W3; a live call confirmed it on 2026-09-26). - .claude/settings.json sets CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION=1500: the 2026-09-26 sweep hit the default 200-per-session cap at 04:10Z, after which Claude workers got no WebSearch results (landscape-sweep-20260926 record, lane limit 10). - blind-adjudicator.md (adoption and examples) and adjudication-prompt.md name astra, gpt-6-sol, gpt-6-luna, gpt-5.6-terra, fable and mythos; adjudicate.py's reported identity words add astra, fable and mythos (the gpt- form already covers the rest). - The harness README documents both budgets. Tests: 312 OK. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ootstrap.md Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
adjudication_provenance() now hashes the changed adjudicate.py, adjudication-prompt.md and blind-adjudicator.md; tests.test_verdict_lane_vendoring requires the current provenance in tools/sota-convergence/lane-provenance.json (CI validate failure on #352). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
force-pushed
the
claude/gpt6-live-search-leakcheck-20260926
branch
from
September 26, 2026 21:33
3e0359a to
23a4e9c
Compare
seathatflowsinourveins
deleted the
claude/gpt6-live-search-leakcheck-20260926
branch
September 26, 2026 21:51
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
…f fit (D1 merit rule) Fold the 2026-09-26 merit directives into Rule 2 of the shared lane prompt, generically and with no repository or product names: "Adoption, incumbency, installation, retained-control status, receipt count, packet position, license, stars and popularity are not evidence of fit; choose among adopted candidates on what their evidence shows was run." License, stars and popularity follow the user's later 2026-09-26 directives: license does not apply to this private experimental environment, and future sessions should judge a repository on its quality and SOTA-ness against the other candidates in its field instead of on unnecessary gates. The prompt keeps exactly five numbered rules, which lane_packets.load_rules parses into every packet's rules field. prompt_sha256 changes from 9b8a8364... to d3ef2cc0.... lane-provenance.json gains entries and changes none: two claude entries (source and vendored workflow paths; the 2026-09-24 workflow, role and transcript-audit hashes with the new prompt hash) and one codex entry (current codex_lane.py with the new prompt hash). No adjudication entry is needed: its prompt is adjudication-prompt.md, which this branch does not change. Rebased onto 59dfe49. Main's lists, #352's adjudication append included, stay an exact prefix and these three entries follow them. Main has registered lane-provenance.json in manifests/evidence.json since this branch's base, so the entry is re-registered under the hot-file protocol; the component-matrix and grand-list generators left their reports unchanged. Tests: Rule 2 must carry the sentence, including in every packet; the Rules section must number exactly 1-5, counted apart from load_rules; the full record-time Claude key, prompt included, must be registered under both workflow paths. The uniqueness test now keys each lane on LANE_PROVENANCE_KEYS: the old Claude tuple left out prompt_sha256, so a prompt-only append collided with the earlier pair (the RR11-2 widening). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
…f fit (D1 merit rule) Fold the 2026-09-26 merit directives into Rule 2 of the shared lane prompt, generically and with no repository or product names: "Adoption, incumbency, installation, retained-control status, receipt count, packet position, license, stars and popularity are not evidence of fit; choose among adopted candidates on what their evidence shows was run." License, stars and popularity follow the user's later 2026-09-26 directives: license does not apply to this private experimental environment, and future sessions should judge a repository on its quality and SOTA-ness against the other candidates in its field instead of on unnecessary gates. The prompt keeps exactly five numbered rules, which lane_packets.load_rules parses into every packet's rules field. prompt_sha256 changes from 9b8a8364... to d3ef2cc0.... lane-provenance.json gains entries and changes none: two claude entries (source and vendored workflow paths; the 2026-09-24 workflow, role and transcript-audit hashes with the new prompt hash) and one codex entry (current codex_lane.py with the new prompt hash). No adjudication entry is needed: its prompt is adjudication-prompt.md, which this branch does not change. Rebased onto 59dfe49. Main's lists, #352's adjudication append included, stay an exact prefix and these three entries follow them. Main has registered lane-provenance.json in manifests/evidence.json since this branch's base, so the entry is re-registered under the hot-file protocol; the component-matrix and grand-list generators left their reports unchanged. Tests: Rule 2 must carry the sentence, including in every packet; the Rules section must number exactly 1-5, counted apart from load_rules; the full record-time Claude key, prompt included, must be registered under both workflow paths. The uniqueness test now keys each lane on LANE_PROVENANCE_KEYS: the old Claude tuple left out prompt_sha256, so a prompt-only append collided with the earlier pair (the RR11-2 widening). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 26, 2026
…f fit (D1 merit rule) (#372) Fold the 2026-09-26 merit directives into Rule 2 of the shared lane prompt, generically and with no repository or product names: "Adoption, incumbency, installation, retained-control status, receipt count, packet position, license, stars and popularity are not evidence of fit; choose among adopted candidates on what their evidence shows was run." License, stars and popularity follow the user's later 2026-09-26 directives: license does not apply to this private experimental environment, and future sessions should judge a repository on its quality and SOTA-ness against the other candidates in its field instead of on unnecessary gates. The prompt keeps exactly five numbered rules, which lane_packets.load_rules parses into every packet's rules field. prompt_sha256 changes from 9b8a8364... to d3ef2cc0.... lane-provenance.json gains entries and changes none: two claude entries (source and vendored workflow paths; the 2026-09-24 workflow, role and transcript-audit hashes with the new prompt hash) and one codex entry (current codex_lane.py with the new prompt hash). No adjudication entry is needed: its prompt is adjudication-prompt.md, which this branch does not change. Rebased onto 59dfe49. Main's lists, #352's adjudication append included, stay an exact prefix and these three entries follow them. Main has registered lane-provenance.json in manifests/evidence.json since this branch's base, so the entry is re-registered under the hot-file protocol; the component-matrix and grand-list generators left their reports unchanged. Tests: Rule 2 must carry the sentence, including in every packet; the Rules section must number exactly 1-5, counted apart from load_rules; the full record-time Claude key, prompt included, must be registered under both workflow paths. The uniqueness test now keys each lane on LANE_PROVENANCE_KEYS: the old Claude tuple left out prompt_sha256, so a prompt-only append collided with the earlier pair (the RR11-2 widening). Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
lane:shared: it changes
.claude/settings.json(project harness config for both lanes) and the blind-adjudication agents. It needs the trading lane's acknowledgement before merge.Three evidence-driven fixes for the verdict wave, found on 2026-09-26.
tools/sota-convergence/landscape-sweep/codex_job.pynow passes-c web_search="live"instead of--search.evidence/artifacts/sota-refresh-20260926/codex/results/websearch-*.json) show that--searchbeforeexec(W1) and no flag (W2) both sendexternal_web_access: false, so search reads a cached index. Onlyweb_search="live"(W3) sendstrue.web_search="live"made a live search on 2026-09-26..claude/settings.jsonsetsCLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION=1500.landscape-sweep-20260926.blind-adjudicator.md(adoption and examples copies) andadjudication-prompt.mdaddastra,gpt-6-sol,gpt-6-luna,gpt-5.6-terra,fableandmythos. Baresol,lunaandterraare ordinary words, so those appear only as slugs.adjudicate.py's reported identity words addastra,fableandmythos; thegpt-form already covers the rest.IdentityWordsTests.The harness README documents both budgets.
adoption/bootstrap.mdmarks the agent change as afterv2026.09.26.2.Tests:
tests.test_adjudicate,test_landscape_sweep_harness,test_install_claude_profileandtest_adoption_statusran 312, OK.tests.test_adoption_docs_consistencypassed 33.validate.pypasses.Rebase (2026-09-26, workstation coordinator). Rebased onto main
015c4702, a second time after #363, #359, #367, #368, #369 and #370 merged, to clear themanifests/evidence.jsonconflict, following the docs/lanes.md hot-file protocol. The manifest is main's copy plus this branch's eleven files: six newly registered (bothblind-adjudicator.mdcopies,test_adjudicate.py,adjudicate.py,adjudication-prompt.md,lane-provenance.json) and five re-hashed (.claude/settings.json,adoption/bootstrap.md,test_landscape_sweep_harness.py, the sweep README andcodex_job.py). Nothing else changed.component_matrix.py --writeandnew_host_grand_list.py --writeleft their reports unchanged. Thelane-provenance.jsonadjudication entry is still appended last, with main's two entries unchanged, and its eight recorded hashes equaladjudicate.adjudication_provenance()on the rebased tree. Local checks at23a4e9c8all exited 0:validate.py,evidence_manifest.py --check,component_matrix.py --check,new_host_grand_list.py --check,build_ecosystem.py --check, and main'sverdict_review_gate.py.test_verdict_lane_vendoring,test_codex_lane,test_lane_packetsandtest_landscape_sweep_harnessran 296, OK (skipped=2). With the 16 other modules that read the changed files, 1366 ran, OK (skipped=21). The frozen sweep templates are unchanged here.Lane consolidation: on 2026-09-26 the user shut down the other lane sessions and assigned this work to the workstation coordinator ("no other peers live, i shut them with the token save repos install into our native workflow and github workflow , harness update and sota foundation of our session as main task"); this quote stands in for the trading lane's ack.
SOTA sources
web_searchconfig (live/cached), measured in Workstation tool refresh: MCP Inspector 2.8.0 and Prometheus 3.15.0 switched, Codex 0.157.1 staged (daemon-safe switch plan) #332'swebsearch-0.155.1.jsonandwebsearch-0.157.1.json.CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION, as read by the installed binary and named in its web-search budget notice. The measured cap hit is inevidence/artifacts/landscape-sweep-20260926/(lane limit 10) once that record merges.🤖 Generated with Claude Code