Repository navigation
Workstation tool refresh: MCP Inspector 2.8.0 and Prometheus 3.15.0 switched, Codex 0.157.1 staged (daemon-safe switch plan) - #332
Merged
Conversation
…facts and receipt) Qualified on nativestack-5975wx-20260925 against the installed 2.7.0: the verified npm tarball (sha256 4db39519..., SLSA provenance from main.yml at refs/tags/2.8.0, commit 1e31c78f) in a new prefix, and 15 gated steps in a bwrap sandbox (private loopback-only network, read-only root, /mnt hidden, MCP_AUTO_OPEN_ENABLED=false) passing 15/15 on both versions with identical outputs: stdio and streamable-HTTP CLI calls against a scratch QMD index. The web auth probe measured PR #2390: with DANGEROUSLY_OMIT_AUTH=false, 0 or yes, 2.7.0 served /api/config without its token and 2.8.0 answered 401. With no Inspector running, bin/mcp-inspector was relinked at 2026-09-26T04:51:23Z and the post-switch CLI check through PATH passed. `mcp-inspector --version` and a bare `mcp-inspector` were never run. - evidence/artifacts/sota-refresh-20260926/mcp-inspector/: harness scripts byte-identical, both arms' results, compare.json, the raw outputs bundled (the harness's token placeholder published as <redacted>), integrity, provenance, package and dependency diffs, relink and post-switch outputs, results.json. - evidence/receipts/mcp-inspector-280-qualification-20260926.json (native_cli_e2e). Not independently reviewed yet. manifests/evidence.json registration is left to the coordinator. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…cts and receipt) Qualified on nativestack-5975wx-20260925 against the running 3.14.0: the v3.15.0 linux-amd64 archive (sha256 2a542df3..., equal to the API digest and sha256sums.txt) installed by the unchanged observability/backends/install.py from a scratch copy with a one-entry pins.json; a side-by-side rehearsal on copies of the live TSDB (same seven targets, 14 rules with health ok, identical six-hour history, SIGHUP reload, no ERROR line) and a rollback check (3.14.0 reopened the TSDB 3.15.0 wrote). The switch installed the unit configure.py renders with the 3.15.0 pin (only the ExecStart prefix differs), restarted ecosystem-prometheus.service at 04:56:49Z (ready in 0.86 s, history before the restart identical) and repointed bin/prometheus and bin/promtool; a 05:17Z read-back found the firing set restored with the original Alertmanager startsAt and no new ntfy message. - evidence/artifacts/sota-refresh-20260926/prometheus/: rehearse.py and switch.py byte-identical, installer pins, checksum and version checks, source checks, the rehearsal results and bundled logs, the unit diff, switch results, stack-command outputs, the read-back, results.json. - evidence/receipts/prometheus-3150-qualification-20260926.json (native_cli_e2e). Not independently reviewed yet. manifests/evidence.json registration is left to the coordinator. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…zed artifacts and receipt) Installed beside 0.155.1 on nativestack-5975wx-20260925 from verified npm tarballs (wrapper 813e2a94..., linux-x64 7f126777..., SLSA provenance from rust-release.yml at refs/tags/rust-v0.157.1; the native binary equals the GitHub release asset's). bin/codex stays on 0.155.1 while the landscape-sweep GPT-6 lanes run it. In a bwrap sandbox (private loopback-only network, scratch CODEX_HOME, no credential read, ~/.codex untouched) and against a loopback fake Responses provider, the sweep lane's argv behaved the same on both versions: events, usage, -o, the output-schema request and the 429 usage-limit events the sweep's limit rule detects. `codex exec --help` is byte-identical; `codex --help` adds only --no-daemon. The exec --json schema adds only an optional web_search `results` field. Finding for the sweep: `codex --search exec` sends external_web_access false (cached mode) on both versions, the same as no flag; only `-c web_search="live"` sends true. A scratch home migrated by 0.157.1 (state_5 55 -> 57, thread_history_1 6 -> 7) still served a 0.155.1 exec turn. - evidence/artifacts/sota-refresh-20260926/codex/: harness and fake provider scripts byte-identical, per-case results, request summaries, bundled events and logs, help outputs, integrity, provenance, binary cross-check, source checks, results.json with the switch and rollback commands for the coordinator. - evidence/receipts/codex-01571-qualification-20260926.json (native_cli_e2e). No pin moves. Not independently reviewed yet. manifests/evidence.json registration is left to the coordinator. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…dex 0.157.1 staged; decision record - manifests/stack.json: mcp-inspector 2.8.0 (its license text now records 2.8.0's MIT-to-Apache-2.0 relicensing statement) and prometheus 3.15.0 (source_pin 5241a27f, the v3.15.0 tag commit), each citing its new receipt. codex stays 0.155.1: 0.157.1 is staged, not switched. - observability/backends/pins.json: prometheus moves to the v3.15.0 linux-amd64 archive (sha256 2a542df3..., checked against the release sha256sums.txt); configure.py renders ExecStart from it. observability/backends/README.md: table, promtool path, release link, and an "Upgrading one backend on an existing host" section with the procedure and rollback this host used. - tests/test_observability_backends_alerts.py: the promtool and amtool paths come from observability/backends/pins.json, so a pin move runs the new binaries (18 tests OK with promtool 3.15.0). - recipes/README.md: the mcp-inspector row names 2.8.0 and the DANGEROUSLY_OMIT_AUTH parse change; the codex row records that a top-level --search does not reach `codex exec` (cached mode on 0.155.1 and 0.157.1; use -c web_search="live") and that 0.157.1 is staged. - blueprints/token-native-focus/saturation-audit.json mirrors the two stack rows. catalogs/landscape/upstream-snapshot.json: fresh gh api rows for mcp-inspector, prometheus and codex; codex joins newer_stable_release_review_queue. - docs/decisions/2026-09-26-workstation-tool-refresh.md: decisions, alternatives and overturn conditions for the three units. catalogs/landscape/foundation.json winner pins do not move (the verdict wave owns them). catalogs/us-equities/agents-operations.* still name Prometheus 3.14.0 and need their own alignment. manifests/evidence.json registration (receipts, artifacts and re-hashed files) is left to the coordinator. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…back (review repair) Review finding: data.rollback only relinked bin/codex and data.switch only checked for 0.155.1 processes, while 0.157.x starts an app-server daemon for interactive launches (daemon_auto_start, stable and on by default at rust-v0.157.1). daemon_probe.py measured it in a sandbox (sandbox_daemon.sh: loopback-only network, empty /run/user, private PID namespace, fresh homes): - a plain interactive 0.157.1 launch copies the 391 MB package into CODEX_HOME/packages/app-server-daemon and starts the app-server and its updater loop from that copy; both outlive the TUI - a plain 0.155.1 launch while that server runs connects to it - `codex app-server daemon stop` leaves the updater loop; SIGTERM ends it - --no-daemon or `codex features disable daemon_auto_start` prevents both; 0.155.1 reads the resulting config.toml (exec turn) and starts its TUI By source (stage/daemon-source-checks.txt, both tags fetched first-hand), the updater runs https://chatgpt.com/codex/install.sh five minutes after start and then hourly to move the copy to the latest release; codex exec has no daemon code. The switch now disables daemon_auto_start before any interactive 0.157.1 launch and checks that no server or package copy exists; the rollback stops the server and its updater loop before relinking. The exact check and stop lines ran in the probe. results.json carries the same commands; the first qualification's run outputs are unchanged. manifests/evidence.json is not edited: the new artifacts need registration by the coordinator. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…hearsal (review repair) Review finding: switch.py logged its verify, daemon-reload and restart exit codes and its post-switch comparisons but repointed bin/prometheus and bin/promtool regardless. In the 2026-09-26 cutover every one of those checks had passed first (cutover/switch.json), so production stands; switch.py and cutover/ stay byte-identical as the as-run record. switch_gated.py moves the links only after every gate passes (preflight, verify, daemon-reload, the loaded unit, restart, readiness, a new MainPID, then version, revision, ExecStart, NRestarts, targets, rules, history and a full-journal ERROR scan). Any failure after the unit install restores the backup unit, with reset-failed before the restart because a crash-looping server exhausts the start limit, and exits nonzero. gate_rehearsal.py ran it on a scratch tmpfs unit with the live unit's restart settings: switch and rollback succeeded; a wrong running version was refused unchanged; a failed daemon-reload, a daemon-reload that silently did not reload (the review's scenario), a new server that crash-looped into the start limit, and a restart that still served 3.14.0 each restored 3.14.0 without moving the links. The production unit and links were read before and after and did not change. The backends README upgrade steps now fail closed. manifests/evidence.json is not edited: the new artifacts need registration. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…refresh Records the review round's two findings (Codex rollback and the background server; the fail-open Prometheus switch script), the background-server behaviour of Codex 0.157.x and the gated Prometheus procedure. The repair has not been re-reviewed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…config with the pin move The adoption template adoption/templates/codex.config.template.toml writes each new host's ~/.codex/config.toml. When the pin moves to 0.157.x, that template needs daemon_auto_start = false under [features], or a new host's first interactive launch starts the self-updating background server outside the pinned prefix. Recorded as an after-switch follow-up in the receipt and results.json; 0.155.1 accepts the key (results/daemon.json T8, T9). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
seathatflowsinourveins
enabled auto-merge (squash)
September 26, 2026 10:40
Owner
Author
|
Trading lane, retroactive check (2026-09-26): no problem for trading. This PR changed only the One knock-on to note:
🤖 Generated with Claude Code |
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
A separate headless Claude session reviewed the 0.157.1 install receipt against the #332 qualification's data.switch and the live host (codex 0.157.1, daemon_auto_start false, no daemon process or package, 0.155.1 kept for rollback) and recorded independent_session agree. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
…t, template, recipes, quota practice evidence/hosts/nativestack-5975wx-20260925/...--codex--install--20260926.json: recorded with scripts/host_receipts.py record from a clean detached worktree at origin/main ca8e37f (published; the catalog_revision), stage install, component codex, version 0.157.1 with --allow-unbound-version because the landscape winner pins stay 0.155.1 until a verdict wave (the qualification receipt's own limitation). Four bounded read-only commands, all exit 0: codex --version (codex-cli 0.157.1), codex features list filtered to daemon_auto_start (stable false), the data.switch daemon process count (0) and package check (absent). It supersedes nothing (no same-day codex install receipt). The quota read is not in it: scripts/codex_quota.py is not on origin/main at that revision. Its manifests/evidence.json files[] entry is left to the coordinator (sha256 b3b8495e..., 5,260 bytes). Pins and manifests, following the qualification receipt's after-switch follow-ups and the #307/#332 pattern: - adoption/pins-linux-x86_64.json: codex 0.157.1, the wrapper tarball URL and sha256 813e2a94..., dist.integrity, the platform package's sha256 and integrity and the daemon note in install_note. - manifests/stack.json: codex 0.157.1, evidence id codex-01571-qualification-20260926 (stack evidence ids must be receipts[] entries; host receipts are files[] only), the release URL and a dated freshness note. - blueprints/token-native-focus/saturation-audit.json: the codex row's version and public receipt. - adoption/templates/codex.config.template.toml: daemon_auto_start = false under [features], so a new host bootstrapped at the 0.157.1 pin does not start the self-updating daemon on its first interactive launch (0.155.1 accepts the key; results/daemon.json). - recipes/README.md: the rust-v0.157.1 codex-package digest (sha256:0e211868..., equal to the GitHub API digest) and the codex row. The macOS pin stays 0.155.1 (its own qualification; no test needs it moved). docs/token-practice.md: the native quota probe with the measured numbers and the shared-budget practice (one host slot pool, the verdict wave first, tell the user at the limit). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
A separate headless Claude session reviewed the 0.157.1 install receipt against the #332 qualification's data.switch and the live host (codex 0.157.1, daemon_auto_start false, no daemon process or package, 0.155.1 kept for rollback) and recorded independent_session agree. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 26, 2026
…a probe and runner gate (#348) * Harden the native Codex quota probe and test it against a protocol fake scripts/codex_quota.py reads the account's usage snapshot through `codex app-server` (stdio JSON-RPC, account/rateLimits/read). Checked against openai/codex codex-rs/app-server-protocol at rust-v0.155.1 and rust-v0.157.1 (common.rs, v1.rs, v2/account.rs, rpc.rs identical for these shapes): messages carry no "jsonrpc" field; initialize sends clientInfo {name, title, version}; the initialized notification follows the initialize answer; the read uses excludeResetCreditDetails (the background-poll form); error answers are {id, error {code, message, data?}}; the top-level rateLimits is the account's "codex" snapshot (app-server account_processor.rs). Hardening over the first draft: - One deadline bounds the whole exchange (selectors on the pipe); the draft's blocking readline never timed out on a silent server. - Cleanup: EOF, then TERM and KILL to the server's own process group, so children that ignore TERM are gone too. - The server runs in an empty temporary directory without RUST_LOG, and a server request is answered with method-not-found. - --gate PERCENT exits 3 when a window's used_percent reaches PERCENT, rateLimitReachedType is set or ordinaryUsageAllowed is false; 2 when no snapshot arrives or nothing can be judged. --json prints exactly one object; accountId and the upsell banner are never printed. tests/test_codex_quota.py (12 tests, no network, no account): snapshot parsing and the exact protocol sequence, interleaved notifications and a server request, the gate outcomes, an error answer, a server that never answers and ignores TERM with a TERM-ignoring child (timeout and group cleanup), a server silent on the read, an early exit, codex missing from PATH. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Landscape sweep: optional quota gate before each GPT-6 job staged.json codex.quota_stop_percent (written by build_args.py --quota-stop-percent; absent, the gate is off) makes the runner, after a job gets its slot and before codex starts, run codex_quota.py --json --gate <percent>. Exit 3 is refused like a usage limit: the job ends with exit 3 before codex starts, <work-dir>/LIMIT (created only when absent, so an earlier reason or a real limit is kept) and the job's stderr.txt name the reason, the used percent and the reset time, and later starts print the marker's reason. Every probe is kept in <job>/quota.json (an attempt file, so it moves to attempts/<n>/ with the rest) and `result` summarizes it as `quota`. A failed probe (no snapshot within codex.quota_timeout_s, default 30 s, an error answer, a missing script) is recorded and never blocks the job. build_args.py stages scripts/codex_quota.py beside the runner as codex_quota.py and records its sha256 under harness.quota_probe; a staged runner uses only that frozen copy, the checkout runner uses scripts/codex_quota.py. The fake codex in the harness tests now answers the app-server quota read; new tests cover the gate off by default, a probe below the stop percent, the refusal and the rerun after a reset, a limit flag below the percent, an existing marker's reason kept, a failed and a missing probe that do not block, bad stop percents and the staged gate. README: files table, a Quota gate coordination entry and the tests note. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Record the Codex 0.157.1 switch: host receipt, Linux pin, stack, audit, template, recipes, quota practice evidence/hosts/nativestack-5975wx-20260925/...--codex--install--20260926.json: recorded with scripts/host_receipts.py record from a clean detached worktree at origin/main ca8e37f (published; the catalog_revision), stage install, component codex, version 0.157.1 with --allow-unbound-version because the landscape winner pins stay 0.155.1 until a verdict wave (the qualification receipt's own limitation). Four bounded read-only commands, all exit 0: codex --version (codex-cli 0.157.1), codex features list filtered to daemon_auto_start (stable false), the data.switch daemon process count (0) and package check (absent). It supersedes nothing (no same-day codex install receipt). The quota read is not in it: scripts/codex_quota.py is not on origin/main at that revision. Its manifests/evidence.json files[] entry is left to the coordinator (sha256 b3b8495e..., 5,260 bytes). Pins and manifests, following the qualification receipt's after-switch follow-ups and the #307/#332 pattern: - adoption/pins-linux-x86_64.json: codex 0.157.1, the wrapper tarball URL and sha256 813e2a94..., dist.integrity, the platform package's sha256 and integrity and the daemon note in install_note. - manifests/stack.json: codex 0.157.1, evidence id codex-01571-qualification-20260926 (stack evidence ids must be receipts[] entries; host receipts are files[] only), the release URL and a dated freshness note. - blueprints/token-native-focus/saturation-audit.json: the codex row's version and public receipt. - adoption/templates/codex.config.template.toml: daemon_auto_start = false under [features], so a new host bootstrapped at the 0.157.1 pin does not start the self-updating daemon on its first interactive launch (0.155.1 accepts the key; results/daemon.json). - recipes/README.md: the rust-v0.157.1 codex-package digest (sha256:0e211868..., equal to the GitHub API digest) and the codex row. The macOS pin stays 0.155.1 (its own qualification; no test needs it moved). docs/token-practice.md: the native quota probe with the measured numbers and the shared-budget practice (one host slot pool, the verdict wave first, tell the user at the limit). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Regenerate the component matrix and new-host grand list for the codex switch python3 scripts/component_matrix.py --write, then python3 scripts/new_host_grand_list.py --write: the new codex install receipt raises the native-clients and agent-sdks codex pass count from 7 to 8 (latest 2026-09-26T14:34:12Z; platform_status unchanged, host_verified), and the Linux bootstrap column reads 0.157.1 while the winner pin and the macOS bootstrap stay 0.155.1. The generators' manifests/evidence.json hash updates are left to the coordinator's registration. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Keep the codex selection consistent: upstream snapshot and the Mac pin lag catalogs/landscape/upstream-snapshot.json: the codex row's selected_version follows manifests/stack.json to 0.157.1 (scripts/ landscape.py requires them equal) and its release_relationship becomes selected_version_matches_latest_stable: the snapshot's own 05:12Z checks already name rust-v0.157.1 (commit 36650394) as the latest stable release and checked its commit. codex leaves the summary's newer-stable review queue, and the recommendation counts four. tests/test_adoption_bootstrap_macos.py: MAC_PIN_LAGS_LINUX records codex 0.155.1 (Mac) against 0.157.1 (Linux) with the qualification receipt, the mechanism #307 used for ai-memory and mcporter. The macOS pin itself stays 0.155.1 until a Mac qualifies 0.157.x. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Label the unretained quota figures in the shared Codex quota note Review finding: the section cited used_percent 61 to 63, the weekly window, reset time and plan, plus a 14:28Z single-bucket 0.77 s read, but no committed receipt or artifact retains that output, and the linked host receipt states the quota read is not part of it. Keep the coordinator-reported figures, label them unretained observations until a sanitized --json read is recorded from a published checkout that contains the script, drop the unretained 14:28Z read, and say the host receipt does not cover the quota read. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * codex_quota: start the app-server with a read-only sandbox The probe only reads account state, so the server should prepare no writable roots (a workspace-write sandbox protects .git mount points inside roots such as /tmp). Live check on codex 0.157.1: the read-only probe returned the snapshot and created nothing under /tmp. Both fakes now require the -c sandbox_mode="read-only" override. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Codex install receipt: independent review (agree) A separate headless Claude session reviewed the 0.157.1 install receipt against the #332 qualification's data.switch and the live host (codex 0.157.1, daemon_auto_start false, no daemon process or package, 0.155.1 kept for rollback) and recorded independent_session agree. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Codex quota probe: repair the GPT-6 review findings (NOT_READY at c0a949f) - codex_quota.py never echoes server error text (backend errors can carry account ids); it reports stage and code with a fixed message. - codex_quota.py stops and reaps the app-server and closes its pipes when setup fails after Popen (an injected EMFILE left them behind). - codex_job.py records a fixed probe-failure message, never the command line (absolute interpreter and work-directory paths), and passes the stop percent and timeout with round-trip precision (95.00001 no longer becomes 95). - docs/token-practice.md quotes no quota figure without a retained, sanitized read. Each finding has a test that fails on c0a949f's code; both suites: 94 tests OK. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Re-register the repaired quota files (hot-file commit) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Mark the Codex pin, recipe row and template as changed after v2026.09.26.2 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Mark the Linux Codex pin change after v2026.09.26.2 on both platform pages Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Re-register after the release markers (hot-file commit) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Sep 26, 2026
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
…vily lead vetting Adds two facts to the record of wf_8397ada1-777 and regenerates it with the packaged tools, as the previous rounds did. - Lane limit 11: every GPT-6-Astra lane, discovery and fit, searched Codex's cached web index, not the live web. The prototype runner started each job as codex --search exec. The #332 Codex qualification (sota-refresh-20260926/codex/results/websearch-0.155.1.json and websearch-0.157.1.json, local stand-in provider) shows --search before exec (W1) and no flag (W2) both sending external_web_access: false, and -c web_search="live" (W3) sending true. The coordinator's observation of a real 0.157.1 live search is not retained. GPT-6 discoveries and votes rested on the cached index, and the release and activity facts GPT-6 cited may lag upstream; with limit 10, no lane after 04:10:13Z had live web search except through gh api and WebFetch. Future runs pass -c web_search="live"; the codex_job.py fix follows in a separate PR. - Regeneration: convert.py with all eleven --limit values (exit 3 for the same two notes; summary identical to the previous round), redaction, id shortening, the credential-name marker and build_manifest.py (pin ids unchanged: 60 occurrences, 39 distinct). returns.json and layers.json are byte-identical; lanes.json and manifest-20260926.json differ only by the added limit. - Ledger: both 2026-09-26 records were re-appended onto origin/main's ledger (unchanged since 3822178). Only manifest_sha256 (both records), one added note on limit 11 (completed record) and the chain hashes changed. - Addendum tavily-leads-vetting.json: the lead vetting of wf_6a6cb7b8-c22 (Sonnet screen, then the sweep's facts refuter and two-family fit refuters). 136 leads in 30 layers, 100 dismissed, 36 proposed, 29 kept, 0 survivors; the Claude fit refuter refuted all 29, GPT-6 21, the facts refuter 4; 8 split cases where GPT-6 did not refute but Claude did. Discovery-completeness evidence, not a verdict, with both search limits (Claude WebSearch capped: 2 calls, both capped; GPT-6 cached). Sanitized with host_receipts.sanitize(); one ref reads <checkout>. Its 23 40-hex ids are shortened to 12 characters, as in returns.json, because the file names sourcegraph/* leads (gitleaks sourcegraph-access-token rule). Its usage record child-usage-wf_6a6cb7b8-c22.json (usage_record.py): complete, 75 children at effort max. - README: GPT-6 cached web search and lead vetting sections, limit 11, files, redaction, verification; limit 11 and the addendum are not yet reviewed. Verification, with a simulated registration of every new and changed file: saturation_ledger.py --check --base origin/main passes (4 records, chain intact) and scripts/validate.py passes. tests.test_saturation_ledger and tests.test_landscape_sweep_harness: 141 run, OK, 2 skipped (no bash 3.2, no shellcheck). tests.test_gitleaks_config and tests.test_verdict_lane_vendoring: 40 run, OK, once another scan's per-user gitleaks lock was free. Without the registration, test_the_committed_ledger_checks fails only on unregistered lane files, as before. manifests/evidence.json is not edited; the coordinator registers the files. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
…ily names in the leak checks - codex_job.py passes -c web_search="live" instead of --search: --search before exec (and no flag) sends external_web_access false (cached index); only web_search="live" sends true (#332 artifacts websearch-*.json, W1-W3; a live call confirmed it on 2026-09-26). - .claude/settings.json sets CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION=1500: the 2026-09-26 sweep hit the default 200-per-session cap at 04:10Z, after which Claude workers got no WebSearch results (landscape-sweep-20260926 record, lane limit 10). - blind-adjudicator.md (adoption and examples) and adjudication-prompt.md name astra, gpt-6-sol, gpt-6-luna, gpt-5.6-terra, fable and mythos; adjudicate.py's reported identity words add astra, fable and mythos (the gpt- form already covers the rest). - The harness README documents both budgets. Tests: 312 OK. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 26, 2026
…est-20260926 for the verdict wave (11 lane limits) (#357) * Landscape-sweep usage: measure runs past 64 KiB, re-run calls and effort deviations Recording the 2026-09-26 sweep (wf_8397ada1-777, 201 agents) exposed three gaps in the packaged harness for a run it legitimately produced: - child-usage.mjs printed its summary and then called process.exit(), which drops stdout writes still pending on a pipe. usage_record.py captures stdout through a pipe, so any run whose summary exceeds 64 KiB was cut at exactly 65,536 bytes and failed to parse (184,831 bytes for this run). It now sets process.exitCode instead, as the Node.js process.exit() documentation advises. - At a Claude usage limit the Workflow paused and re-ran its 8 waiting agents after the reset under the same journal key ("Re-running 8 waiting agents"). child-usage.mjs counted each earlier attempt as an incomplete child, so a run in which every call returned read as incomplete, and convert.py measured those votes from the failed attempt ("<synthetic>+claude-opus-5-5"). child-usage.mjs now lists an attempt that returned nothing and whose key started again under superseded_attempts (with superseded_by), keeps its issues and effort check, and still counts its usage in by_resolved_model; convert.py takes the attempt that returned and never names <synthetic>. - A worker measured at another effort than max (here a pinned skill with `effort: low` frontmatter lowered the final turn of three workers) made make_result.py refuse the whole record. convert.py --usage now records each such worker as an effort_deviation retained failure of its layer (the critic's of every layer), which reopens the layer; make_result.py accepts the usage only when every deviation is recorded so. Tests: new cases in test-child-usage.mjs and tests/test_landscape_sweep_harness.py fail on 3822178 and pass here; SHA256SUMS updated for the two changed files. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Landscape-sweep source reviews: review a Hugging Face model survivor at its Hub commit The 2026-09-26 sweep (wf_8397ada1-777) kept a Hugging Face model repository in agents-models-workers (https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B), a legitimate survivor of a model layer. source_reviews.py read only GitHub through gh api, refused it as "not a GitHub repository URL", and make_result.py then refused the record for a survivor without a source review. source_reviews.py now reviews a Hugging Face model repository at the commit of its default revision. The Hub's model-info endpoint /api/models/<repo_id> (the endpoint huggingface_hub's HfApi.model_info calls, checked at huggingface_hub main 62a1f2fa8383) gives the commit, the card license and the repository state. The model card is read at that commit through the documented "Resolve a file" endpoint (https://huggingface.co/.well-known/openapi.json), anonymously, and its YAML metadata block is not excerpted. The review is named hf-<namespace>-<name>.json. A Hub dataset or Space URL is still reported and skipped. The GitHub path is unchanged apart from sharing the excerpt helper. Test: SourceReviewTests.test_a_hugging_face_model_survivor_is_reviewed_at_its_hub_commit (stubbed Hub fetch, no network) fails before this change and passes with it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Saturation ledger test: derive the seed-only state from the two seed records test_no_seed_layer_counts_as_clean asserted that no layer of the whole committed ledger counts as clean. That held while the ledger held only the 2026-09-23 seed records; a later completed sweep with clean layers (landscape-sweep-20260926 leaves agent-sdks, backtesting-engine and token-efficiency at 1) is its own evidence and made the whole-ledger assertion false. The test now derives the state from the two seed records, which is what its name and the seed checks around it assert. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Landscape-sweep source reviews: repairs from the independent review An independent headless review (claude -p, Opus, read-only) of the Hugging Face support found three low-severity defects. Each is now covered by SourceReviewTests.test_hub_reviews_survive_truncated_responses_trailing_slashes_and_indented_card_metadata, which fails on the previous code and passes here: - hub_get caught URLError, OSError and ValueError, but http.client.HTTPException (IncompleteRead, BadStatusLine) is not an OSError. One truncated Hub response therefore aborted the whole run with a traceback before any review was written. It is now a HubError and is reported and skipped like any other. - Survivors were grouped by sweep_common.slug(), which keeps a Hugging Face URL whole, so the same model with and without a trailing slash got two reviews. repository_key() now groups a model as hf:<namespace>/<name>. make_result.py matches a review to a survivor as the ledger compares repositories (no trailing slash, lowercased), so it finds the review either way. - The card's YAML metadata block was stripped only when the card began with "---". It now uses huggingface_hub's repocard.REGEX_YAML_BLOCK (checked at main 62a1f2fa8383), which allows leading whitespace. A model-info response that is not a JSON object is also refused as a HubError. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * gitleaks: exempt the reviewed pin commit ids of manifest-20260926 by value catalogs/sota-convergence/manifest-20260926.json (landscape-sweep-20260926) names the rule id "sourcegraph-access-token" once, in a secrets-credentials refuter's reasoning. The rule's keyword then makes every 40-hex value in the file a candidate. The lane's own free-text commit ids are shortened to 12 characters in the retained evidence. The only 40-hex values left are the 39 distinct git commit ids of the 60 pins that the generator copies from the catalog layer files, and those pins come in free-form shapes such as "v0.44.0; source <id>" and "<id> (blueprints/...)". The allowlist is rule-scoped and limited to that exact file. It exempts only a finding whose secret is exactly one of those reviewed ids, pinned by value like the ai-memory fingerprints. It does not exempt a line shape: an independent review of a first, line-shaped version found that any bare 40-hex value on a pin line, such as a legacy Sourcegraph token or an uppercase variant, would have passed. An unreviewed 40-hex value on a pin line or elsewhere, the uppercase form of a reviewed id, an sgp_-prefixed token, and the reviewed ids in any other file all stay detected. Tests: test_d3 and test_d4 in tests/test_gitleaks_config.py read the pinned ids from .gitleaks.toml, because a 40-hex literal in the test file would itself be a finding. Both fail without the allowlist and pass with it. The full suite passes with GITLEAKS_TESTS_REQUIRED=1 on the pinned gitleaks 8.30.1 (27 tests, none skipped). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Landscape sweep 2026-09-26: record wf_8397ada1-777 and its stopped first attempt Records the 2026-09-26 saturation sweep of all 32 layers (Workflow run wf_8397ada1-777, 201 Claude children at effort max, 80 GPT-6-Astra jobs) with the packaged tools/sota-convergence/landscape-sweep/ harness: - usage_record.py: child-usage complete, with 201 children and 8 superseded attempts (re-run after a Claude usage-limit pause). Three effort deviations (the property-based-testing skill's effort: low) are recorded as retained failures. - convert.py: 32 layers, 287 proposals and 56 survivors (54 repositories), with no lost or excluded rounds. 7 layers have Claude-only discovery, 9 layers have no first-round GPT-6 fit votes, and 18 layers are reopened. GPT-6 jobs: ok 64, failed_exit_null 8, failed_exit_3 6, failed_exit_1 2; copy check match 64. Nine run-specific lane limits cover the prototype harness, the cut known-repository slugs, the GPT-6 failures and their causes (the job timeline is retained in gpt6-jobs.json), the seed mapping and the 12 unsurfaced seeds, effort, the Claude usage-limit re-run, call scope, the Codex review step and the earlier attempts. - Redactions: host strings with host_receipts.sanitize(); the manifest builder's credential-name marker; free-text 40-hex commit ids shortened to 12 characters, because the file names the sourcegraph-access-token rule. - source_reviews.py: one upstream-provenance review per surviving repository (53 GitHub, 1 Hugging Face model). - build_manifest.py: catalogs/sota-convergence/manifest-20260926.json. Its Codex review step is this run's GPT-6 lanes, not a separate review. - make_result.py and saturation_ledger.py --append: landscape-sweep-20260926 (completed) after landscape-sweep-20260926-attempt-1 (wf_a874897e-af1, stopped, lower-bound usage). Reopen entries also copy the saturation report's current pin_moved and selection_changed triggers. The one-layer smoke wf_1753e674-5dc is retained as an attempt without a ledger record. - The Tavily Research cross-check is kept as discovery leads only (no report text, hashed request ids). - independent-review.json records a separate headless Opus review (read-only): needs_changes with six low findings, all repaired in this branch (one repair round). Evidence class: proposals, refuter votes and upstream source reviews. Nothing was installed and no winner changed. manifests/evidence.json is not edited here: the coordinator registers the new and changed files. With a simulated registration, saturation_ledger.py --check --base origin/main, scripts/validate.py and the harness and ledger tests pass. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Landscape-sweep tools: record capped WebSearch calls and refutations by absence Repairs two defects that the review of the 2026-09-26 record found in the packaged lane. WebSearch session cap. Claude Code allows one session at most CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION WebSearch calls (default 200 since v2.1.212), counted across the main conversation and every subagent. A capped call returns a notice that tells the worker to go on without searching (tools-reference "Session search limit"; env-vars; the 2.1.283 client's default is Q=200). The lane's budgets allow 480 discovery and 640 refuter searches, so a full sweep can pass the cap unnoticed: - child-usage.mjs counts each child's WebSearch calls and capped calls (web_search, from tool_use/tool_result pairs; the notice must open the result section) and totals them per run. Neither usage nor exit codes change. - convert.py --usage makes every capped worker a web_search_capped retained failure of its layer (the critic's of every layer), as with effort_deviation. - make_result.py refuses a usage record without a measured web_search, or one whose capped workers the returns do not list as that failure. - The harness README (Coordination), the recipe and usage_record.py say how to raise the cap before a full sweep. Refuted by absence. A proposal refuted only because a vote did not return (the GPT-6 fit vote in nine layers of 2026-09-26) counted downstream as refuted on merit: as known in later sweeps and as previous_sweep.refuted in the next discovery input. - convert.py marks a missing facts vote {missing: true} (fit members already were) and names such proposals in the layer's votes_note. - saturation_ledger.py: refutes_on_merit and refuted_by_absence read the marker; adjudicated_repos leaves those entries out in --check and --append, so a later re-proposal stays new. Existing records are unaffected: none before 2026-09-26 has such an entry. - build_inputs.py lists them under previous_sweep.not_adjudicated with a note. Tests: node test-child-usage.mjs (46), tests.test_landscape_sweep_harness and tests.test_saturation_ledger (one expected failure until the coordinator registers this branch's evidence), tests.test_verdict_lane_vendoring; SHA256SUMS updated for the two changed example files. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Landscape sweep 2026-09-26 record: WebSearch cap and refutations by absence The repair round for the record review's two findings, rebuilt with the tools of 1bd2416b. - Usage: the three usage records were re-measured with child-usage.mjs at 1bd2416b. Every usage figure is unchanged, and each record gains web_search. In wf_8397ada1-777, 200 WebSearch calls were made, 154 returned results and 46 were capped, in 30 workers. The first capped call was at 04:10:13Z and the last at 11:55:57Z. Run 1 made 25 calls and the smoke 1, none capped. The compact attempt records copy those counts. - convert.py: counts unchanged (32 layers, 287 proposals, 56 survivors, no lost or excluded round; copy check match 64). returns.json differs from the previous record only by 61 added web_search_capped failures: 29 for workers in 16 layers and one for the critic in each of the 32 layers. All 32 layers are reopened. No layer is clean. agent-sdks, backtesting-engine and token-efficiency were clean before this repair. - Lane limits: limit 3 now states the 27 proposals refuted by absence, limit 7 gives the measured WebSearch counts, and the new limit 10 covers the session cap. The manifest was rebuilt; only its lane limits changed, and its 39 pin ids are unchanged. - Ledger: both 2026-09-26 records were re-appended onto origin/main's ledger. The records gain nine votes_note entries, the new reopen entries and a note on the cap. known/new, survived and refuted are unchanged. - README: new sections on the WebSearch session cap and on refuted by absence, a regenerated layer table, and the second review round with the rejected alternative (new GPT-6 fit votes would come from a new model run, not from this run). independent-review.json records that second round. Verification, with a simulated registration of every new and changed file: saturation_ledger.py --check --base origin/main passes, and scripts/validate.py passes. The CI-style suites pass: test_saturation_ledger, test_landscape_sweep_harness, test_verdict_lane_vendoring and test_gitleaks_config (181 tests, 2 skipped for bash 3.2 and shellcheck). node test-child-usage.mjs passes (46). The coordinator registers the files in manifests/evidence.json. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Landscape sweep 2026-09-26 record: GPT-6 cached web search and the Tavily lead vetting Adds two facts to the record of wf_8397ada1-777 and regenerates it with the packaged tools, as the previous rounds did. - Lane limit 11: every GPT-6-Astra lane, discovery and fit, searched Codex's cached web index, not the live web. The prototype runner started each job as codex --search exec. The #332 Codex qualification (sota-refresh-20260926/codex/results/websearch-0.155.1.json and websearch-0.157.1.json, local stand-in provider) shows --search before exec (W1) and no flag (W2) both sending external_web_access: false, and -c web_search="live" (W3) sending true. The coordinator's observation of a real 0.157.1 live search is not retained. GPT-6 discoveries and votes rested on the cached index, and the release and activity facts GPT-6 cited may lag upstream; with limit 10, no lane after 04:10:13Z had live web search except through gh api and WebFetch. Future runs pass -c web_search="live"; the codex_job.py fix follows in a separate PR. - Regeneration: convert.py with all eleven --limit values (exit 3 for the same two notes; summary identical to the previous round), redaction, id shortening, the credential-name marker and build_manifest.py (pin ids unchanged: 60 occurrences, 39 distinct). returns.json and layers.json are byte-identical; lanes.json and manifest-20260926.json differ only by the added limit. - Ledger: both 2026-09-26 records were re-appended onto origin/main's ledger (unchanged since 3822178). Only manifest_sha256 (both records), one added note on limit 11 (completed record) and the chain hashes changed. - Addendum tavily-leads-vetting.json: the lead vetting of wf_6a6cb7b8-c22 (Sonnet screen, then the sweep's facts refuter and two-family fit refuters). 136 leads in 30 layers, 100 dismissed, 36 proposed, 29 kept, 0 survivors; the Claude fit refuter refuted all 29, GPT-6 21, the facts refuter 4; 8 split cases where GPT-6 did not refute but Claude did. Discovery-completeness evidence, not a verdict, with both search limits (Claude WebSearch capped: 2 calls, both capped; GPT-6 cached). Sanitized with host_receipts.sanitize(); one ref reads <checkout>. Its 23 40-hex ids are shortened to 12 characters, as in returns.json, because the file names sourcegraph/* leads (gitleaks sourcegraph-access-token rule). Its usage record child-usage-wf_6a6cb7b8-c22.json (usage_record.py): complete, 75 children at effort max. - README: GPT-6 cached web search and lead vetting sections, limit 11, files, redaction, verification; limit 11 and the addendum are not yet reviewed. Verification, with a simulated registration of every new and changed file: saturation_ledger.py --check --base origin/main passes (4 records, chain intact) and scripts/validate.py passes. tests.test_saturation_ledger and tests.test_landscape_sweep_harness: 141 run, OK, 2 skipped (no bash 3.2, no shellcheck). tests.test_gitleaks_config and tests.test_verdict_lane_vendoring: 40 run, OK, once another scan's per-user gitleaks lock was free. Without the registration, test_the_committed_ledger_checks fails only on unregistered lane files, as before. manifests/evidence.json is not edited; the coordinator registers the files. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Register the landscape-sweep-20260926 record (hot-file commit) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Landscape-sweep tools: superseded attempts keep usage-integrity failures; failure coverage per worker and layer Repairs the two P2 findings of a read-only GPT-6-Astra review (max effort, live search) of the 2026-09-26 record's checkout. Each repair comes with a regression test that failed before it. (A) child-usage.mjs: a superseded attempt whose transcript held an assistant message without provider usage left the run status "complete", although by_resolved_model could not count that usage, so downstream usage could be marked complete falsely. An attempt's usage-integrity failures (an assistant message without provider usage or without a resolved model, or no transcript file) are now its usage_issues; a superseded attempt with any makes the run incomplete (the reason names them) and the CLI exit 1. Its other issues (no result entry, the <synthetic> usage-limit row outside the family) stay expected, and an attempt that made no request still counts as zero usage. usage_record.py's summary names such attempts (superseded_usage_issues). test-child-usage.mjs: four new checks failed before the repair (missing usage, unresolved model, no transcript, CLI exit); a guard for an attempt without requests passes. SHA256SUMS updated for the two changed files. (B) make_result.py: failure coverage was checked over all layers at once, so removing one layer's critic-cap failure and its reopen entry still passed (another layer's "critic" failure satisfied the check) and that layer derived a clean count of 1. effort_deviation and web_search_capped failures are now checked per worker and per layer, with the mapping convert.py records them by (deviation_rounds, moved unchanged to sweep_common.py): a <role>:<layer>[:followup] worker's failure in that layer's round, the critic's in every layer of the record; a worker naming no layer of the record is refused as before. test_failure_coverage_is_checked_per_worker_and_per_layer drops the critic's cap failure, then its effort deviation, from one of two healthy layers, and moves a worker's failure to the other layer; all three subtests failed before the repair. Docs: the harness README (evidence contract, WebSearch cap, review repairs) and the workflows README. Tests: test-child-usage.mjs 51 passed, test-usage-receipts.mjs 18, test-envelope.mjs 214, test-contract-mutations.mjs 47; test-codex-envelope.mjs and check-syntax.mjs exit 0. tests.test_saturation_ledger, tests.test_landscape_sweep_harness, tests.test_adoption_docs_consistency and tests.test_verdict_lane_vendoring: 198 run, OK (3 skipped: no bash 3.2, no shellcheck, no profile coverage change since the pinned release). manifests/evidence.json is not edited; the coordinator registers the changed files. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Landscape sweep 2026-09-26 record: re-measured usage, GPT-6 review round, lane limit 12 Regenerated once with the pipeline of the earlier rounds, on the tools of 7760b1da, so the manifest sha changes once. - Usage: the three usage records were re-measured with child-usage.mjs at 7760b1da (sha256 85b95813...730d). Each child_usage is identical to the 1bd2416b measurement: wf_8397ada1-777 stays complete (none of its 8 superseded attempts holds uncounted usage) and run 1 stays incomplete, so lower_bound_usage and every ledger outcome are unchanged. Only tool_commit, tool_sha256 and measured_at_utc changed. - Lane limit 12, a dated correction from the trading lane's review of #357: the execution-broker proposal of wboayue/rust-ibapi (a re-pin through nautilus_trader PR #5041) attributes #4983 to rc5. Checked with gh api: #4983's body reports Version v1.227.0; 1b0a49d2 is identical to tag v2.0.0rc5, and there crates/adapters/interactive_brokers/src/execution/ core.rs:493-495 overrides handles_order_venue to return true, so the ClientVenueMismatch denial behind crates/execution/src/engine/mod.rs:2222 cannot fire for the IB client; the open PR #280 reclassifies #4983 in catalogs/us-equities/runtime-target.json. Read instead: rc5 IBKR execution is unqualified (no rc5 stock order observed). The proposal text is kept. - convert.py with limits 1-12: exit 3 on the same two notes, summary identical. returns.json and layers.json are byte-identical; lanes.json and manifest-20260926.json differ only by the added limit (pin ids unchanged: 60 occurrences, 39 distinct). - make_result.py with the per-worker, per-layer coverage check passes on this record and gives the same RESULT.json. - Ledger: both 2026-09-26 records were re-appended onto origin/main's ledger (803bc35, unchanged since 3822178). Only manifest_sha256, usage_sha256 and the chain hashes changed. - README and independent-review.json: the third review round (GPT-6-Astra, NOT_READY, two P2 findings, both reproduced and repaired in 7760b1da) and the trading lane's correction. Verification, with a simulated registration of the 18 changed registered files: saturation_ledger.py --check --base origin/main passes (4 records, chain intact) and scripts/validate.py passes. Tests: tests.test_saturation_ledger, tests.test_landscape_sweep_harness, tests.test_adoption_docs_consistency and tests.test_verdict_lane_vendoring 198 run, OK (3 skipped: no bash 3.2, no shellcheck, no profile coverage change); test-child-usage.mjs 51 passed; tests.test_gitleaks_config 27 OK. manifests/evidence.json is not edited; the coordinator registers the changed files. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Re-register the repaired sweep record (hot-file commit) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
…ily names in the leak checks - codex_job.py passes -c web_search="live" instead of --search: --search before exec (and no flag) sends external_web_access false (cached index); only web_search="live" sends true (#332 artifacts websearch-*.json, W1-W3; a live call confirmed it on 2026-09-26). - .claude/settings.json sets CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION=1500: the 2026-09-26 sweep hit the default 200-per-session cap at 04:10Z, after which Claude workers got no WebSearch results (landscape-sweep-20260926 record, lane limit 10). - blind-adjudicator.md (adoption and examples) and adjudication-prompt.md name astra, gpt-6-sol, gpt-6-luna, gpt-5.6-terra, fable and mythos; adjudicate.py's reported identity words add astra, fable and mythos (the gpt- form already covers the rest). - The harness README documents both budgets. Tests: 312 OK. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
…ily names in the leak checks - codex_job.py passes -c web_search="live" instead of --search: --search before exec (and no flag) sends external_web_access false (cached index); only web_search="live" sends true (#332 artifacts websearch-*.json, W1-W3; a live call confirmed it on 2026-09-26). - .claude/settings.json sets CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION=1500: the 2026-09-26 sweep hit the default 200-per-session cap at 04:10Z, after which Claude workers got no WebSearch results (landscape-sweep-20260926 record, lane limit 10). - blind-adjudicator.md (adoption and examples) and adjudication-prompt.md name astra, gpt-6-sol, gpt-6-luna, gpt-5.6-terra, fable and mythos; adjudicate.py's reported identity words add astra, fable and mythos (the gpt- form already covers the rest). - The harness README documents both budgets. Tests: 312 OK. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 26, 2026
…ude family names in the blind-adjudication leak checks (#352) * GPT-6 live web search, WebSearch session budget, and GPT-6/Claude family names in the leak checks - codex_job.py passes -c web_search="live" instead of --search: --search before exec (and no flag) sends external_web_access false (cached index); only web_search="live" sends true (#332 artifacts websearch-*.json, W1-W3; a live call confirmed it on 2026-09-26). - .claude/settings.json sets CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION=1500: the 2026-09-26 sweep hit the default 200-per-session cap at 04:10Z, after which Claude workers got no WebSearch results (landscape-sweep-20260926 record, lane limit 10). - blind-adjudicator.md (adoption and examples) and adjudication-prompt.md name astra, gpt-6-sol, gpt-6-luna, gpt-5.6-terra, fable and mythos; adjudicate.py's reported identity words add astra, fable and mythos (the gpt- form already covers the rest). - The harness README documents both budgets. Tests: 312 OK. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Mark the blind-adjudicator leak-check change after v2026.09.26.2 in bootstrap.md Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Register the new adjudication provenance for the leak-check change adjudication_provenance() now hashes the changed adjudicate.py, adjudication-prompt.md and blind-adjudicator.md; tests.test_verdict_lane_vendoring requires the current provenance in tools/sota-convergence/lane-provenance.json (CI validate failure on #352). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Re-register after the provenance entry (hot-file commit) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR refreshes the workstation tools that had drifted from upstream and that no other session owns. It follows the 2026-09-25 refresh pattern (#291/#307): each new version is installed into a new versioned prefix beside the old one, qualified against it, switched with the old prefix kept for rollback, and recorded with a receipt. Winner pins in
catalogs/landscape/*.jsonare unchanged; the verdict wave owns them.dist.integrity;npm audit signaturesverified 190 signatures and 34 attestations; SLSA provenancemain.yml@refs/tags/2.8.0. In a bwrap sandbox (loopback-only network, no browser), both versions passed 15/15 identical CLI steps. The PR #2390 fix was measured: withDANGEROUSLY_OMIT_AUTHset tofalse,0oryes, 2.7.0 served/api/configwithout a token and 2.8.0 returned 401. Note: the licence statement moves from MIT towards Apache-2.0, recorded instack.jsonecosystem-prometheus.serviceready 0.86 s later; 3.14.0 keptsha256sums.txt. A rehearsal on copies of the live TSDB showed the same 7 targets, 14 rules healthy and identical 6-hour history. The rollback probe passed (3.14.0 reopens the TSDB 3.15.0 wrote). After the switch, history read back identical and firing alerts kept theirstartsAt. Newswitch_gated.py, rehearsed on a scratch unitbin/codexstays on 0.155.1 until the sweep's GPT-6 lanes finishdist.integrity, with signatures and attestations verified (rust-release.yml@refs/tags/rust-v0.157.1). The native binary is byte-identical to the GitHub asset. Against a fake provider, the lane'sexec --jsonargv behaves the same, including the 429 usage-limit events; only an optionalweb_search.resultsfield is addedCodex 0.157.1 caution, found in review and measured in a sandbox. The first interactive launch installs a background app-server daemon: a ~391 MB package copy under
CODEX_HOME/packages. Its updater (read from source, not measured) runshttps://chatgpt.com/codex/install.shfive minutes after start and then hourly. A 0.155.1 launch then connects to that daemon, so relinking alone would not roll back. The receipt'sdata.switchtherefore:codex features disable daemon_auto_startwith the staged binary before any interactive 0.157.1 launch;data.rollbackincludes the TERM loop that also ends the updater;daemon stopalone left it running.Review record
Tests
Full suite CI-style with registration simulated: 5,805 tests, 648 skipped, 1 failure. The failure is a known timing flake,
test_adaptive_paper_metricsFileSd, which passes 4/4 in isolation and is not caused by this branch.tests.test_observability_backends_alertsran 18 OK with promtool 3.15.0. After registration on current main,validate.pypassed (157 receipts).SOTA sources
@modelcontextprotocol/inspector@2.8.0(dist.integrity, provenance) and modelcontextprotocol/inspector PR #2390.sha256sums.txtand tag commit 5241a27f.@openai/codex@0.157.1tarballs; daemon behaviour read from source at both tags and measured in a sandbox.docs/decisions/2026-09-26-workstation-tool-refresh.mdhas the full record.🤖 Generated with Claude Code