Skip to content

Workstation tool refresh: MCP Inspector 2.8.0 and Prometheus 3.15.0 switched, Codex 0.157.1 staged (daemon-safe switch plan) - #332

Merged
seathatflowsinourveins merged 9 commits into
mainfrom
claude/tool-refresh-20260926
Sep 26, 2026
Merged

seathatflowsinourveins merged 9 commits into
mainfrom
claude/tool-refresh-20260926

Conversation

@seathatflowsinourveins

Copy link
Copy Markdown
Owner

This PR refreshes the workstation tools that had drifted from upstream and that no other session owns. It follows the 2026-09-25 refresh pattern (#291/#307): each new version is installed into a new versioned prefix beside the old one, qualified against it, switched with the old prefix kept for rollback, and recorded with a receipt. Winner pins in catalogs/landscape/*.json are unchanged; the verdict wave owns them.

Tool Outcome Key evidence
MCP Inspector 2.7.0 → 2.8.0 Switched 04:51Z; 2.7.0 prefix kept The tarball sha512 matches npm dist.integrity; npm audit signatures verified 190 signatures and 34 attestations; SLSA provenance main.yml@refs/tags/2.8.0. In a bwrap sandbox (loopback-only network, no browser), both versions passed 15/15 identical CLI steps. The PR #2390 fix was measured: with DANGEROUSLY_OMIT_AUTH set to false, 0 or yes, 2.7.0 served /api/config without a token and 2.8.0 returned 401. Note: the licence statement moves from MIT towards Apache-2.0, recorded in stack.json
Prometheus 3.14.0 → 3.15.0 Switched 04:56:49Z; ecosystem-prometheus.service ready 0.86 s later; 3.14.0 kept Archive sha256 matches sha256sums.txt. A rehearsal on copies of the live TSDB showed the same 7 targets, 14 rules healthy and identical 6-hour history. The rollback probe passed (3.14.0 reopens the TSDB 3.15.0 wrote). After the switch, history read back identical and firing alerts kept their startsAt. New switch_gated.py, rehearsed on a scratch unit
Codex CLI 0.155.1 → 0.157.1 Staged only; bin/codex stays on 0.155.1 until the sweep's GPT-6 lanes finish Both tarballs match dist.integrity, with signatures and attestations verified (rust-release.yml@refs/tags/rust-v0.157.1). The native binary is byte-identical to the GitHub asset. Against a fake provider, the lane's exec --json argv behaves the same, including the 429 usage-limit events; only an optional web_search.results field is added

Codex 0.157.1 caution, found in review and measured in a sandbox. The first interactive launch installs a background app-server daemon: a ~391 MB package copy under CODEX_HOME/packages. Its updater (read from source, not measured) runs https://chatgpt.com/codex/install.sh five minutes after start and then hourly. A 0.155.1 launch then connects to that daemon, so relinking alone would not roll back. The receipt's data.switch therefore:

  1. runs codex features disable daemon_auto_start with the staged binary before any interactive 0.157.1 launch;
  2. reads the setting back;
  3. checks that no daemon process or package copy exists after the relink and again after the first launch.

data.rollback includes the TERM loop that also ends the updater; daemon stop alone left it running.

Review record

  • Claude Opus: 1 medium (the rollback ignored the daemon) and 4 low.
  • GPT-6-Astra: 1 medium (the Prometheus switch gate).
  • Both mediums were confirmed and fixed in one repair round (9258c37c, a7c8694c, cdd42501). The repair has not been re-reviewed, and both receipts say so.

Tests

Full suite CI-style with registration simulated: 5,805 tests, 648 skipped, 1 failure. The failure is a known timing flake, test_adaptive_paper_metrics FileSd, which passes 4/4 in isolation and is not caused by this branch. tests.test_observability_backends_alerts ran 18 OK with promtool 3.15.0. After registration on current main, validate.py passed (157 receipts).

SOTA sources

🤖 Generated with Claude Code

Scout and others added 9 commits September 26, 2026 06:39
…facts and receipt)

Qualified on nativestack-5975wx-20260925 against the installed 2.7.0:
the verified npm tarball (sha256 4db39519..., SLSA provenance from
main.yml at refs/tags/2.8.0, commit 1e31c78f) in a new prefix, and 15
gated steps in a bwrap sandbox (private loopback-only network,
read-only root, /mnt hidden, MCP_AUTO_OPEN_ENABLED=false) passing 15/15
on both versions with identical outputs: stdio and streamable-HTTP CLI
calls against a scratch QMD index. The web auth probe measured PR #2390:
with DANGEROUSLY_OMIT_AUTH=false, 0 or yes, 2.7.0 served /api/config
without its token and 2.8.0 answered 401. With no Inspector running,
bin/mcp-inspector was relinked at 2026-09-26T04:51:23Z and the
post-switch CLI check through PATH passed. `mcp-inspector --version`
and a bare `mcp-inspector` were never run.

- evidence/artifacts/sota-refresh-20260926/mcp-inspector/: harness
  scripts byte-identical, both arms' results, compare.json, the raw
  outputs bundled (the harness's token placeholder published as
  <redacted>), integrity, provenance, package and dependency diffs,
  relink and post-switch outputs, results.json.
- evidence/receipts/mcp-inspector-280-qualification-20260926.json
  (native_cli_e2e).

Not independently reviewed yet. manifests/evidence.json registration is
left to the coordinator.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…cts and receipt)

Qualified on nativestack-5975wx-20260925 against the running 3.14.0:
the v3.15.0 linux-amd64 archive (sha256 2a542df3..., equal to the API
digest and sha256sums.txt) installed by the unchanged
observability/backends/install.py from a scratch copy with a one-entry
pins.json; a side-by-side rehearsal on copies of the live TSDB (same
seven targets, 14 rules with health ok, identical six-hour history,
SIGHUP reload, no ERROR line) and a rollback check (3.14.0 reopened the
TSDB 3.15.0 wrote). The switch installed the unit configure.py renders
with the 3.15.0 pin (only the ExecStart prefix differs), restarted
ecosystem-prometheus.service at 04:56:49Z (ready in 0.86 s, history
before the restart identical) and repointed bin/prometheus and
bin/promtool; a 05:17Z read-back found the firing set restored with the
original Alertmanager startsAt and no new ntfy message.

- evidence/artifacts/sota-refresh-20260926/prometheus/: rehearse.py and
  switch.py byte-identical, installer pins, checksum and version checks,
  source checks, the rehearsal results and bundled logs, the unit diff,
  switch results, stack-command outputs, the read-back, results.json.
- evidence/receipts/prometheus-3150-qualification-20260926.json
  (native_cli_e2e).

Not independently reviewed yet. manifests/evidence.json registration is
left to the coordinator.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…zed artifacts and receipt)

Installed beside 0.155.1 on nativestack-5975wx-20260925 from verified
npm tarballs (wrapper 813e2a94..., linux-x64 7f126777..., SLSA
provenance from rust-release.yml at refs/tags/rust-v0.157.1; the native
binary equals the GitHub release asset's). bin/codex stays on 0.155.1
while the landscape-sweep GPT-6 lanes run it.

In a bwrap sandbox (private loopback-only network, scratch CODEX_HOME,
no credential read, ~/.codex untouched) and against a loopback fake
Responses provider, the sweep lane's argv behaved the same on both
versions: events, usage, -o, the output-schema request and the 429
usage-limit events the sweep's limit rule detects. `codex exec --help`
is byte-identical; `codex --help` adds only --no-daemon. The exec --json
schema adds only an optional web_search `results` field.

Finding for the sweep: `codex --search exec` sends
external_web_access false (cached mode) on both versions, the same as no
flag; only `-c web_search="live"` sends true. A scratch home migrated by
0.157.1 (state_5 55 -> 57, thread_history_1 6 -> 7) still served a
0.155.1 exec turn.

- evidence/artifacts/sota-refresh-20260926/codex/: harness and fake
  provider scripts byte-identical, per-case results, request summaries,
  bundled events and logs, help outputs, integrity, provenance, binary
  cross-check, source checks, results.json with the switch and rollback
  commands for the coordinator.
- evidence/receipts/codex-01571-qualification-20260926.json
  (native_cli_e2e).

No pin moves. Not independently reviewed yet. manifests/evidence.json
registration is left to the coordinator.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…dex 0.157.1 staged; decision record

- manifests/stack.json: mcp-inspector 2.8.0 (its license text now records
  2.8.0's MIT-to-Apache-2.0 relicensing statement) and prometheus 3.15.0
  (source_pin 5241a27f, the v3.15.0 tag commit), each citing its new
  receipt. codex stays 0.155.1: 0.157.1 is staged, not switched.
- observability/backends/pins.json: prometheus moves to the v3.15.0
  linux-amd64 archive (sha256 2a542df3..., checked against the release
  sha256sums.txt); configure.py renders ExecStart from it.
  observability/backends/README.md: table, promtool path, release link,
  and an "Upgrading one backend on an existing host" section with the
  procedure and rollback this host used.
- tests/test_observability_backends_alerts.py: the promtool and amtool
  paths come from observability/backends/pins.json, so a pin move runs
  the new binaries (18 tests OK with promtool 3.15.0).
- recipes/README.md: the mcp-inspector row names 2.8.0 and the
  DANGEROUSLY_OMIT_AUTH parse change; the codex row records that a
  top-level --search does not reach `codex exec` (cached mode on 0.155.1
  and 0.157.1; use -c web_search="live") and that 0.157.1 is staged.
- blueprints/token-native-focus/saturation-audit.json mirrors the two
  stack rows. catalogs/landscape/upstream-snapshot.json: fresh gh api
  rows for mcp-inspector, prometheus and codex; codex joins
  newer_stable_release_review_queue.
- docs/decisions/2026-09-26-workstation-tool-refresh.md: decisions,
  alternatives and overturn conditions for the three units.

catalogs/landscape/foundation.json winner pins do not move (the verdict
wave owns them). catalogs/us-equities/agents-operations.* still name
Prometheus 3.14.0 and need their own alignment. manifests/evidence.json
registration (receipts, artifacts and re-hashed files) is left to the
coordinator.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…back (review repair)

Review finding: data.rollback only relinked bin/codex and data.switch only
checked for 0.155.1 processes, while 0.157.x starts an app-server daemon for
interactive launches (daemon_auto_start, stable and on by default at
rust-v0.157.1).

daemon_probe.py measured it in a sandbox (sandbox_daemon.sh: loopback-only
network, empty /run/user, private PID namespace, fresh homes):
- a plain interactive 0.157.1 launch copies the 391 MB package into
  CODEX_HOME/packages/app-server-daemon and starts the app-server and its
  updater loop from that copy; both outlive the TUI
- a plain 0.155.1 launch while that server runs connects to it
- `codex app-server daemon stop` leaves the updater loop; SIGTERM ends it
- --no-daemon or `codex features disable daemon_auto_start` prevents both;
  0.155.1 reads the resulting config.toml (exec turn) and starts its TUI
By source (stage/daemon-source-checks.txt, both tags fetched first-hand),
the updater runs https://chatgpt.com/codex/install.sh five minutes after
start and then hourly to move the copy to the latest release; codex exec
has no daemon code.

The switch now disables daemon_auto_start before any interactive 0.157.1
launch and checks that no server or package copy exists; the rollback stops
the server and its updater loop before relinking. The exact check and stop
lines ran in the probe. results.json carries the same commands; the first
qualification's run outputs are unchanged. manifests/evidence.json is not
edited: the new artifacts need registration by the coordinator.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…hearsal (review repair)

Review finding: switch.py logged its verify, daemon-reload and restart exit
codes and its post-switch comparisons but repointed bin/prometheus and
bin/promtool regardless. In the 2026-09-26 cutover every one of those checks
had passed first (cutover/switch.json), so production stands; switch.py and
cutover/ stay byte-identical as the as-run record.

switch_gated.py moves the links only after every gate passes (preflight,
verify, daemon-reload, the loaded unit, restart, readiness, a new MainPID,
then version, revision, ExecStart, NRestarts, targets, rules, history and a
full-journal ERROR scan). Any failure after the unit install restores the
backup unit, with reset-failed before the restart because a crash-looping
server exhausts the start limit, and exits nonzero.

gate_rehearsal.py ran it on a scratch tmpfs unit with the live unit's
restart settings: switch and rollback succeeded; a wrong running version was
refused unchanged; a failed daemon-reload, a daemon-reload that silently did
not reload (the review's scenario), a new server that crash-looped into the
start limit, and a restart that still served 3.14.0 each restored 3.14.0
without moving the links. The production unit and links were read before and
after and did not change. The backends README upgrade steps now fail closed.

manifests/evidence.json is not edited: the new artifacts need registration.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…refresh

Records the review round's two findings (Codex rollback and the background
server; the fail-open Prometheus switch script), the background-server
behaviour of Codex 0.157.x and the gated Prometheus procedure. The repair
has not been re-reviewed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…config with the pin move

The adoption template adoption/templates/codex.config.template.toml writes
each new host's ~/.codex/config.toml. When the pin moves to 0.157.x, that
template needs daemon_auto_start = false under [features], or a new host's
first interactive launch starts the self-updating background server outside
the pinned prefix. Recorded as an after-switch follow-up in the receipt and
results.json; 0.155.1 accepts the key (results/daemon.json T8, T9).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@seathatflowsinourveins
seathatflowsinourveins enabled auto-merge (squash) September 26, 2026 10:40
@seathatflowsinourveins
seathatflowsinourveins merged commit 2e77144 into main Sep 26, 2026
25 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the claude/tool-refresh-20260926 branch September 26, 2026 10:58
@seathatflowsinourveins seathatflowsinourveins added the lane:shared Touches files owned by both lanes; needs both lanes' acknowledgement label Sep 26, 2026
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Trading lane, retroactive check (2026-09-26): no problem for trading.

This PR changed only the mcp-inspector and prometheus components in manifests/stack.json, and no trading path.

One knock-on to note: tests/test_observability_backends_alerts.py hard-codes ecosystem-prometheus-3.14.0/promtool, including the #284 rule unit tests for EquitiesPaperMetricsMissing.

  • On a host that keeps only 3.15.0, those promtool tests would skip rather than run.
  • Deriving the tool path from observability/backends/pins.json would keep them native after the bump.

🤖 Generated with Claude Code

seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
A separate headless Claude session reviewed the 0.157.1 install receipt against the
#332 qualification's data.switch and the live host (codex 0.157.1, daemon_auto_start
false, no daemon process or package, 0.155.1 kept for rollback) and recorded
independent_session agree.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
…t, template, recipes, quota practice

evidence/hosts/nativestack-5975wx-20260925/...--codex--install--20260926.json:
recorded with scripts/host_receipts.py record from a clean detached
worktree at origin/main ca8e37f (published; the catalog_revision),
stage install, component codex, version 0.157.1 with
--allow-unbound-version because the landscape winner pins stay 0.155.1
until a verdict wave (the qualification receipt's own limitation).
Four bounded read-only commands, all exit 0: codex --version
(codex-cli 0.157.1), codex features list filtered to daemon_auto_start
(stable false), the data.switch daemon process count (0) and package
check (absent). It supersedes nothing (no same-day codex install
receipt). The quota read is not in it: scripts/codex_quota.py is not on
origin/main at that revision. Its manifests/evidence.json files[] entry
is left to the coordinator (sha256 b3b8495e..., 5,260 bytes).

Pins and manifests, following the qualification receipt's after-switch
follow-ups and the #307/#332 pattern:
- adoption/pins-linux-x86_64.json: codex 0.157.1, the wrapper tarball
  URL and sha256 813e2a94..., dist.integrity, the platform package's
  sha256 and integrity and the daemon note in install_note.
- manifests/stack.json: codex 0.157.1, evidence id
  codex-01571-qualification-20260926 (stack evidence ids must be
  receipts[] entries; host receipts are files[] only), the release URL
  and a dated freshness note.
- blueprints/token-native-focus/saturation-audit.json: the codex row's
  version and public receipt.
- adoption/templates/codex.config.template.toml: daemon_auto_start =
  false under [features], so a new host bootstrapped at the 0.157.1 pin
  does not start the self-updating daemon on its first interactive
  launch (0.155.1 accepts the key; results/daemon.json).
- recipes/README.md: the rust-v0.157.1 codex-package digest
  (sha256:0e211868..., equal to the GitHub API digest) and the codex row.
The macOS pin stays 0.155.1 (its own qualification; no test needs it
moved).

docs/token-practice.md: the native quota probe with the measured
numbers and the shared-budget practice (one host slot pool, the verdict
wave first, tell the user at the limit).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
A separate headless Claude session reviewed the 0.157.1 install receipt against the
#332 qualification's data.switch and the live host (codex 0.157.1, daemon_auto_start
false, no daemon process or package, 0.155.1 kept for rollback) and recorded
independent_session agree.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 26, 2026
…a probe and runner gate (#348)

* Harden the native Codex quota probe and test it against a protocol fake

scripts/codex_quota.py reads the account's usage snapshot through
`codex app-server` (stdio JSON-RPC, account/rateLimits/read). Checked
against openai/codex codex-rs/app-server-protocol at rust-v0.155.1 and
rust-v0.157.1 (common.rs, v1.rs, v2/account.rs, rpc.rs identical for
these shapes): messages carry no "jsonrpc" field; initialize sends
clientInfo {name, title, version}; the initialized notification follows
the initialize answer; the read uses excludeResetCreditDetails (the
background-poll form); error answers are {id, error {code, message,
data?}}; the top-level rateLimits is the account's "codex" snapshot
(app-server account_processor.rs).

Hardening over the first draft:
- One deadline bounds the whole exchange (selectors on the pipe); the
  draft's blocking readline never timed out on a silent server.
- Cleanup: EOF, then TERM and KILL to the server's own process group,
  so children that ignore TERM are gone too.
- The server runs in an empty temporary directory without RUST_LOG, and
  a server request is answered with method-not-found.
- --gate PERCENT exits 3 when a window's used_percent reaches PERCENT,
  rateLimitReachedType is set or ordinaryUsageAllowed is false; 2 when
  no snapshot arrives or nothing can be judged. --json prints exactly
  one object; accountId and the upsell banner are never printed.

tests/test_codex_quota.py (12 tests, no network, no account): snapshot
parsing and the exact protocol sequence, interleaved notifications and
a server request, the gate outcomes, an error answer, a server that
never answers and ignores TERM with a TERM-ignoring child (timeout and
group cleanup), a server silent on the read, an early exit, codex
missing from PATH.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Landscape sweep: optional quota gate before each GPT-6 job

staged.json codex.quota_stop_percent (written by build_args.py
--quota-stop-percent; absent, the gate is off) makes the runner, after a
job gets its slot and before codex starts, run codex_quota.py --json
--gate <percent>. Exit 3 is refused like a usage limit: the job ends
with exit 3 before codex starts, <work-dir>/LIMIT (created only when
absent, so an earlier reason or a real limit is kept) and the job's
stderr.txt name the reason, the used percent and the reset time, and
later starts print the marker's reason. Every probe is kept in
<job>/quota.json (an attempt file, so it moves to attempts/<n>/ with
the rest) and `result` summarizes it as `quota`. A failed probe (no
snapshot within codex.quota_timeout_s, default 30 s, an error answer, a
missing script) is recorded and never blocks the job.

build_args.py stages scripts/codex_quota.py beside the runner as
codex_quota.py and records its sha256 under harness.quota_probe; a
staged runner uses only that frozen copy, the checkout runner uses
scripts/codex_quota.py. The fake codex in the harness tests now answers
the app-server quota read; new tests cover the gate off by default, a
probe below the stop percent, the refusal and the rerun after a reset,
a limit flag below the percent, an existing marker's reason kept, a
failed and a missing probe that do not block, bad stop percents and the
staged gate. README: files table, a Quota gate coordination entry and
the tests note.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Record the Codex 0.157.1 switch: host receipt, Linux pin, stack, audit, template, recipes, quota practice

evidence/hosts/nativestack-5975wx-20260925/...--codex--install--20260926.json:
recorded with scripts/host_receipts.py record from a clean detached
worktree at origin/main ca8e37f (published; the catalog_revision),
stage install, component codex, version 0.157.1 with
--allow-unbound-version because the landscape winner pins stay 0.155.1
until a verdict wave (the qualification receipt's own limitation).
Four bounded read-only commands, all exit 0: codex --version
(codex-cli 0.157.1), codex features list filtered to daemon_auto_start
(stable false), the data.switch daemon process count (0) and package
check (absent). It supersedes nothing (no same-day codex install
receipt). The quota read is not in it: scripts/codex_quota.py is not on
origin/main at that revision. Its manifests/evidence.json files[] entry
is left to the coordinator (sha256 b3b8495e..., 5,260 bytes).

Pins and manifests, following the qualification receipt's after-switch
follow-ups and the #307/#332 pattern:
- adoption/pins-linux-x86_64.json: codex 0.157.1, the wrapper tarball
  URL and sha256 813e2a94..., dist.integrity, the platform package's
  sha256 and integrity and the daemon note in install_note.
- manifests/stack.json: codex 0.157.1, evidence id
  codex-01571-qualification-20260926 (stack evidence ids must be
  receipts[] entries; host receipts are files[] only), the release URL
  and a dated freshness note.
- blueprints/token-native-focus/saturation-audit.json: the codex row's
  version and public receipt.
- adoption/templates/codex.config.template.toml: daemon_auto_start =
  false under [features], so a new host bootstrapped at the 0.157.1 pin
  does not start the self-updating daemon on its first interactive
  launch (0.155.1 accepts the key; results/daemon.json).
- recipes/README.md: the rust-v0.157.1 codex-package digest
  (sha256:0e211868..., equal to the GitHub API digest) and the codex row.
The macOS pin stays 0.155.1 (its own qualification; no test needs it
moved).

docs/token-practice.md: the native quota probe with the measured
numbers and the shared-budget practice (one host slot pool, the verdict
wave first, tell the user at the limit).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Regenerate the component matrix and new-host grand list for the codex switch

python3 scripts/component_matrix.py --write, then
python3 scripts/new_host_grand_list.py --write: the new codex install
receipt raises the native-clients and agent-sdks codex pass count from
7 to 8 (latest 2026-09-26T14:34:12Z; platform_status unchanged,
host_verified), and the Linux bootstrap column reads 0.157.1 while the
winner pin and the macOS bootstrap stay 0.155.1. The generators'
manifests/evidence.json hash updates are left to the coordinator's
registration.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Keep the codex selection consistent: upstream snapshot and the Mac pin lag

catalogs/landscape/upstream-snapshot.json: the codex row's
selected_version follows manifests/stack.json to 0.157.1 (scripts/
landscape.py requires them equal) and its release_relationship becomes
selected_version_matches_latest_stable: the snapshot's own 05:12Z
checks already name rust-v0.157.1 (commit 36650394) as the latest
stable release and checked its commit. codex leaves the summary's
newer-stable review queue, and the recommendation counts four.

tests/test_adoption_bootstrap_macos.py: MAC_PIN_LAGS_LINUX records
codex 0.155.1 (Mac) against 0.157.1 (Linux) with the qualification
receipt, the mechanism #307 used for ai-memory and mcporter. The macOS
pin itself stays 0.155.1 until a Mac qualifies 0.157.x.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Label the unretained quota figures in the shared Codex quota note

Review finding: the section cited used_percent 61 to 63, the weekly window,
reset time and plan, plus a 14:28Z single-bucket 0.77 s read, but no committed
receipt or artifact retains that output, and the linked host receipt states the
quota read is not part of it. Keep the coordinator-reported figures, label them
unretained observations until a sanitized --json read is recorded from a
published checkout that contains the script, drop the unretained 14:28Z read,
and say the host receipt does not cover the quota read.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* codex_quota: start the app-server with a read-only sandbox

The probe only reads account state, so the server should prepare no writable roots (a
workspace-write sandbox protects .git mount points inside roots such as /tmp). Live
check on codex 0.157.1: the read-only probe returned the snapshot and created nothing
under /tmp. Both fakes now require the -c sandbox_mode="read-only" override.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Codex install receipt: independent review (agree)

A separate headless Claude session reviewed the 0.157.1 install receipt against the
#332 qualification's data.switch and the live host (codex 0.157.1, daemon_auto_start
false, no daemon process or package, 0.155.1 kept for rollback) and recorded
independent_session agree.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Codex quota probe: repair the GPT-6 review findings (NOT_READY at c0a949f)

- codex_quota.py never echoes server error text (backend errors can carry account ids);
  it reports stage and code with a fixed message.
- codex_quota.py stops and reaps the app-server and closes its pipes when setup fails
  after Popen (an injected EMFILE left them behind).
- codex_job.py records a fixed probe-failure message, never the command line (absolute
  interpreter and work-directory paths), and passes the stop percent and timeout with
  round-trip precision (95.00001 no longer becomes 95).
- docs/token-practice.md quotes no quota figure without a retained, sanitized read.
Each finding has a test that fails on c0a949f's code; both suites: 94 tests OK.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Re-register the repaired quota files (hot-file commit)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Mark the Codex pin, recipe row and template as changed after v2026.09.26.2

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Mark the Linux Codex pin change after v2026.09.26.2 on both platform pages

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Re-register after the release markers (hot-file commit)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
…vily lead vetting

Adds two facts to the record of wf_8397ada1-777 and regenerates it with the
packaged tools, as the previous rounds did.

- Lane limit 11: every GPT-6-Astra lane, discovery and fit, searched Codex's
  cached web index, not the live web. The prototype runner started each job
  as codex --search exec. The #332 Codex qualification
  (sota-refresh-20260926/codex/results/websearch-0.155.1.json and
  websearch-0.157.1.json, local stand-in provider) shows --search before exec
  (W1) and no flag (W2) both sending external_web_access: false, and
  -c web_search="live" (W3) sending true. The coordinator's observation of a
  real 0.157.1 live search is not retained. GPT-6 discoveries and votes rested
  on the cached index, and the release and activity facts GPT-6 cited may lag
  upstream; with limit 10, no lane after 04:10:13Z had live web search except
  through gh api and WebFetch. Future runs pass -c web_search="live"; the
  codex_job.py fix follows in a separate PR.
- Regeneration: convert.py with all eleven --limit values (exit 3 for the same
  two notes; summary identical to the previous round), redaction, id
  shortening, the credential-name marker and build_manifest.py (pin ids
  unchanged: 60 occurrences, 39 distinct). returns.json and layers.json are
  byte-identical; lanes.json and manifest-20260926.json differ only by the
  added limit.
- Ledger: both 2026-09-26 records were re-appended onto origin/main's ledger
  (unchanged since 3822178). Only manifest_sha256 (both records), one added
  note on limit 11 (completed record) and the chain hashes changed.
- Addendum tavily-leads-vetting.json: the lead vetting of wf_6a6cb7b8-c22
  (Sonnet screen, then the sweep's facts refuter and two-family fit
  refuters). 136 leads in 30 layers, 100 dismissed, 36 proposed, 29 kept,
  0 survivors; the Claude fit refuter refuted all 29, GPT-6 21, the facts
  refuter 4; 8 split cases where GPT-6 did not refute but Claude did.
  Discovery-completeness evidence, not a verdict, with both search limits
  (Claude WebSearch capped: 2 calls, both capped; GPT-6 cached). Sanitized
  with host_receipts.sanitize(); one ref reads <checkout>. Its 23 40-hex ids
  are shortened to 12 characters, as in returns.json, because the file names
  sourcegraph/* leads (gitleaks sourcegraph-access-token rule). Its usage
  record child-usage-wf_6a6cb7b8-c22.json (usage_record.py): complete, 75
  children at effort max.
- README: GPT-6 cached web search and lead vetting sections, limit 11, files,
  redaction, verification; limit 11 and the addendum are not yet reviewed.

Verification, with a simulated registration of every new and changed file:
saturation_ledger.py --check --base origin/main passes (4 records, chain
intact) and scripts/validate.py passes. tests.test_saturation_ledger and
tests.test_landscape_sweep_harness: 141 run, OK, 2 skipped (no bash 3.2, no
shellcheck). tests.test_gitleaks_config and tests.test_verdict_lane_vendoring:
40 run, OK, once another scan's per-user gitleaks lock was free. Without the
registration, test_the_committed_ledger_checks fails only on unregistered
lane files, as before. manifests/evidence.json is not edited; the coordinator
registers the files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
…ily names in the leak checks

- codex_job.py passes -c web_search="live" instead of --search: --search before exec (and
  no flag) sends external_web_access false (cached index); only web_search="live" sends
  true (#332 artifacts websearch-*.json, W1-W3; a live call confirmed it on 2026-09-26).
- .claude/settings.json sets CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION=1500: the 2026-09-26
  sweep hit the default 200-per-session cap at 04:10Z, after which Claude workers got no
  WebSearch results (landscape-sweep-20260926 record, lane limit 10).
- blind-adjudicator.md (adoption and examples) and adjudication-prompt.md name astra,
  gpt-6-sol, gpt-6-luna, gpt-5.6-terra, fable and mythos; adjudicate.py's reported
  identity words add astra, fable and mythos (the gpt- form already covers the rest).
- The harness README documents both budgets. Tests: 312 OK.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 26, 2026
…est-20260926 for the verdict wave (11 lane limits) (#357)

* Landscape-sweep usage: measure runs past 64 KiB, re-run calls and effort deviations

Recording the 2026-09-26 sweep (wf_8397ada1-777, 201 agents) exposed three gaps
in the packaged harness for a run it legitimately produced:

- child-usage.mjs printed its summary and then called process.exit(), which
  drops stdout writes still pending on a pipe. usage_record.py captures stdout
  through a pipe, so any run whose summary exceeds 64 KiB was cut at exactly
  65,536 bytes and failed to parse (184,831 bytes for this run). It now sets
  process.exitCode instead, as the Node.js process.exit() documentation advises.
- At a Claude usage limit the Workflow paused and re-ran its 8 waiting agents
  after the reset under the same journal key ("Re-running 8 waiting agents").
  child-usage.mjs counted each earlier attempt as an incomplete child, so a run
  in which every call returned read as incomplete, and convert.py measured
  those votes from the failed attempt ("<synthetic>+claude-opus-5-5").
  child-usage.mjs now lists an attempt that returned nothing and whose key
  started again under superseded_attempts (with superseded_by), keeps its
  issues and effort check, and still counts its usage in by_resolved_model;
  convert.py takes the attempt that returned and never names <synthetic>.
- A worker measured at another effort than max (here a pinned skill with
  `effort: low` frontmatter lowered the final turn of three workers) made
  make_result.py refuse the whole record. convert.py --usage now records each
  such worker as an effort_deviation retained failure of its layer (the
  critic's of every layer), which reopens the layer; make_result.py accepts
  the usage only when every deviation is recorded so.

Tests: new cases in test-child-usage.mjs and tests/test_landscape_sweep_harness.py
fail on 3822178 and pass here; SHA256SUMS updated for the two changed files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Landscape-sweep source reviews: review a Hugging Face model survivor at its Hub commit

The 2026-09-26 sweep (wf_8397ada1-777) kept a Hugging Face model repository in
agents-models-workers (https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B), a
legitimate survivor of a model layer. source_reviews.py read only GitHub through
gh api, refused it as "not a GitHub repository URL", and make_result.py then
refused the record for a survivor without a source review.

source_reviews.py now reviews a Hugging Face model repository at the commit of
its default revision. The Hub's model-info endpoint /api/models/<repo_id> (the
endpoint huggingface_hub's HfApi.model_info calls, checked at huggingface_hub
main 62a1f2fa8383) gives the commit, the card license and the repository state.
The model card is read at that commit through the documented "Resolve a file"
endpoint (https://huggingface.co/.well-known/openapi.json), anonymously, and its
YAML metadata block is not excerpted. The review is named
hf-<namespace>-<name>.json. A Hub dataset or Space URL is still reported and
skipped. The GitHub path is unchanged apart from sharing the excerpt helper.

Test: SourceReviewTests.test_a_hugging_face_model_survivor_is_reviewed_at_its_hub_commit
(stubbed Hub fetch, no network) fails before this change and passes with it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Saturation ledger test: derive the seed-only state from the two seed records

test_no_seed_layer_counts_as_clean asserted that no layer of the whole committed
ledger counts as clean. That held while the ledger held only the 2026-09-23 seed
records; a later completed sweep with clean layers (landscape-sweep-20260926
leaves agent-sdks, backtesting-engine and token-efficiency at 1) is its own
evidence and made the whole-ledger assertion false. The test now derives the
state from the two seed records, which is what its name and the seed checks
around it assert.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Landscape-sweep source reviews: repairs from the independent review

An independent headless review (claude -p, Opus, read-only) of the Hugging Face
support found three low-severity defects. Each is now covered by
SourceReviewTests.test_hub_reviews_survive_truncated_responses_trailing_slashes_and_indented_card_metadata,
which fails on the previous code and passes here:

- hub_get caught URLError, OSError and ValueError, but http.client.HTTPException
  (IncompleteRead, BadStatusLine) is not an OSError. One truncated Hub response
  therefore aborted the whole run with a traceback before any review was
  written. It is now a HubError and is reported and skipped like any other.
- Survivors were grouped by sweep_common.slug(), which keeps a Hugging Face URL
  whole, so the same model with and without a trailing slash got two reviews.
  repository_key() now groups a model as hf:<namespace>/<name>. make_result.py
  matches a review to a survivor as the ledger compares repositories (no
  trailing slash, lowercased), so it finds the review either way.
- The card's YAML metadata block was stripped only when the card began with
  "---". It now uses huggingface_hub's repocard.REGEX_YAML_BLOCK (checked at
  main 62a1f2fa8383), which allows leading whitespace.

A model-info response that is not a JSON object is also refused as a HubError.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* gitleaks: exempt the reviewed pin commit ids of manifest-20260926 by value

catalogs/sota-convergence/manifest-20260926.json (landscape-sweep-20260926)
names the rule id "sourcegraph-access-token" once, in a secrets-credentials
refuter's reasoning. The rule's keyword then makes every 40-hex value in the
file a candidate. The lane's own free-text commit ids are shortened to 12
characters in the retained evidence. The only 40-hex values left are the 39
distinct git commit ids of the 60 pins that the generator copies from the
catalog layer files, and those pins come in free-form shapes such as
"v0.44.0; source <id>" and "<id> (blueprints/...)".

The allowlist is rule-scoped and limited to that exact file. It exempts only
a finding whose secret is exactly one of those reviewed ids, pinned by value
like the ai-memory fingerprints. It does not exempt a line shape: an
independent review of a first, line-shaped version found that any bare 40-hex
value on a pin line, such as a legacy Sourcegraph token or an uppercase
variant, would have passed. An unreviewed 40-hex value on a pin line or
elsewhere, the uppercase form of a reviewed id, an sgp_-prefixed token, and the
reviewed ids in any other file all stay detected.

Tests: test_d3 and test_d4 in tests/test_gitleaks_config.py read the pinned ids
from .gitleaks.toml, because a 40-hex literal in the test file would itself be
a finding. Both fail without the allowlist and pass with it. The full suite
passes with GITLEAKS_TESTS_REQUIRED=1 on the pinned gitleaks 8.30.1 (27 tests,
none skipped).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Landscape sweep 2026-09-26: record wf_8397ada1-777 and its stopped first attempt

Records the 2026-09-26 saturation sweep of all 32 layers (Workflow run
wf_8397ada1-777, 201 Claude children at effort max, 80 GPT-6-Astra jobs) with
the packaged tools/sota-convergence/landscape-sweep/ harness:

- usage_record.py: child-usage complete, with 201 children and 8 superseded
  attempts (re-run after a Claude usage-limit pause). Three effort deviations
  (the property-based-testing skill's effort: low) are recorded as retained
  failures.
- convert.py: 32 layers, 287 proposals and 56 survivors (54 repositories), with
  no lost or excluded rounds. 7 layers have Claude-only discovery, 9 layers have
  no first-round GPT-6 fit votes, and 18 layers are reopened. GPT-6 jobs: ok 64,
  failed_exit_null 8, failed_exit_3 6, failed_exit_1 2; copy check match 64.
  Nine run-specific lane limits cover the prototype harness, the cut
  known-repository slugs, the GPT-6 failures and their causes (the job timeline
  is retained in gpt6-jobs.json), the seed mapping and the 12 unsurfaced
  seeds, effort, the Claude usage-limit re-run, call scope, the Codex review
  step and the earlier attempts.
- Redactions: host strings with host_receipts.sanitize(); the manifest
  builder's credential-name marker; free-text 40-hex commit ids shortened to
  12 characters, because the file names the sourcegraph-access-token rule.
- source_reviews.py: one upstream-provenance review per surviving repository
  (53 GitHub, 1 Hugging Face model).
- build_manifest.py: catalogs/sota-convergence/manifest-20260926.json. Its
  Codex review step is this run's GPT-6 lanes, not a separate review.
- make_result.py and saturation_ledger.py --append: landscape-sweep-20260926
  (completed) after landscape-sweep-20260926-attempt-1 (wf_a874897e-af1,
  stopped, lower-bound usage). Reopen entries also copy the saturation
  report's current pin_moved and selection_changed triggers. The one-layer
  smoke wf_1753e674-5dc is retained as an attempt without a ledger record.
- The Tavily Research cross-check is kept as discovery leads only (no report
  text, hashed request ids).
- independent-review.json records a separate headless Opus review (read-only):
  needs_changes with six low findings, all repaired in this branch (one repair
  round).

Evidence class: proposals, refuter votes and upstream source reviews. Nothing
was installed and no winner changed. manifests/evidence.json is not edited here:
the coordinator registers the new and changed files. With a simulated
registration, saturation_ledger.py --check --base origin/main, scripts/validate.py
and the harness and ledger tests pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Landscape-sweep tools: record capped WebSearch calls and refutations by absence

Repairs two defects that the review of the 2026-09-26 record found in the packaged lane.

WebSearch session cap. Claude Code allows one session at most
CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION WebSearch calls (default 200 since
v2.1.212), counted across the main conversation and every subagent. A capped
call returns a notice that tells the worker to go on without searching
(tools-reference "Session search limit"; env-vars; the 2.1.283 client's
default is Q=200). The lane's budgets allow 480 discovery and 640 refuter
searches, so a full sweep can pass the cap unnoticed:
- child-usage.mjs counts each child's WebSearch calls and capped calls
  (web_search, from tool_use/tool_result pairs; the notice must open the
  result section) and totals them per run. Neither usage nor exit codes change.
- convert.py --usage makes every capped worker a web_search_capped retained
  failure of its layer (the critic's of every layer), as with effort_deviation.
- make_result.py refuses a usage record without a measured web_search, or one
  whose capped workers the returns do not list as that failure.
- The harness README (Coordination), the recipe and usage_record.py say how to
  raise the cap before a full sweep.

Refuted by absence. A proposal refuted only because a vote did not return
(the GPT-6 fit vote in nine layers of 2026-09-26) counted downstream as
refuted on merit: as known in later sweeps and as previous_sweep.refuted in
the next discovery input.
- convert.py marks a missing facts vote {missing: true} (fit members already
  were) and names such proposals in the layer's votes_note.
- saturation_ledger.py: refutes_on_merit and refuted_by_absence read the
  marker; adjudicated_repos leaves those entries out in --check and --append,
  so a later re-proposal stays new. Existing records are unaffected: none
  before 2026-09-26 has such an entry.
- build_inputs.py lists them under previous_sweep.not_adjudicated with a note.

Tests: node test-child-usage.mjs (46), tests.test_landscape_sweep_harness and
tests.test_saturation_ledger (one expected failure until the coordinator
registers this branch's evidence), tests.test_verdict_lane_vendoring;
SHA256SUMS updated for the two changed example files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Landscape sweep 2026-09-26 record: WebSearch cap and refutations by absence

The repair round for the record review's two findings, rebuilt with the tools
of 1bd2416b.

- Usage: the three usage records were re-measured with child-usage.mjs at
  1bd2416b. Every usage figure is unchanged, and each record gains web_search.
  In wf_8397ada1-777, 200 WebSearch calls were made, 154 returned results and
  46 were capped, in 30 workers. The first capped call was at 04:10:13Z and
  the last at 11:55:57Z. Run 1 made 25 calls and the smoke 1, none capped.
  The compact attempt records copy those counts.
- convert.py: counts unchanged (32 layers, 287 proposals, 56 survivors, no
  lost or excluded round; copy check match 64). returns.json differs from the
  previous record only by 61 added web_search_capped failures: 29 for workers
  in 16 layers and one for the critic in each of the 32 layers. All 32 layers
  are reopened. No layer is clean. agent-sdks, backtesting-engine and
  token-efficiency were clean before this repair.
- Lane limits: limit 3 now states the 27 proposals refuted by absence, limit
  7 gives the measured WebSearch counts, and the new limit 10 covers the
  session cap. The manifest was rebuilt; only its lane limits changed, and
  its 39 pin ids are unchanged.
- Ledger: both 2026-09-26 records were re-appended onto origin/main's ledger.
  The records gain nine votes_note entries, the new reopen entries and a
  note on the cap. known/new, survived and refuted are unchanged.
- README: new sections on the WebSearch session cap and on refuted by
  absence, a regenerated layer table, and the second review round with the
  rejected alternative (new GPT-6 fit votes would come from a new model run,
  not from this run). independent-review.json records that second round.

Verification, with a simulated registration of every new and changed file:
saturation_ledger.py --check --base origin/main passes, and scripts/validate.py
passes. The CI-style suites pass: test_saturation_ledger,
test_landscape_sweep_harness, test_verdict_lane_vendoring and
test_gitleaks_config (181 tests, 2 skipped for bash 3.2 and shellcheck).
node test-child-usage.mjs passes (46). The coordinator registers the files in
manifests/evidence.json.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Landscape sweep 2026-09-26 record: GPT-6 cached web search and the Tavily lead vetting

Adds two facts to the record of wf_8397ada1-777 and regenerates it with the
packaged tools, as the previous rounds did.

- Lane limit 11: every GPT-6-Astra lane, discovery and fit, searched Codex's
  cached web index, not the live web. The prototype runner started each job
  as codex --search exec. The #332 Codex qualification
  (sota-refresh-20260926/codex/results/websearch-0.155.1.json and
  websearch-0.157.1.json, local stand-in provider) shows --search before exec
  (W1) and no flag (W2) both sending external_web_access: false, and
  -c web_search="live" (W3) sending true. The coordinator's observation of a
  real 0.157.1 live search is not retained. GPT-6 discoveries and votes rested
  on the cached index, and the release and activity facts GPT-6 cited may lag
  upstream; with limit 10, no lane after 04:10:13Z had live web search except
  through gh api and WebFetch. Future runs pass -c web_search="live"; the
  codex_job.py fix follows in a separate PR.
- Regeneration: convert.py with all eleven --limit values (exit 3 for the same
  two notes; summary identical to the previous round), redaction, id
  shortening, the credential-name marker and build_manifest.py (pin ids
  unchanged: 60 occurrences, 39 distinct). returns.json and layers.json are
  byte-identical; lanes.json and manifest-20260926.json differ only by the
  added limit.
- Ledger: both 2026-09-26 records were re-appended onto origin/main's ledger
  (unchanged since 3822178). Only manifest_sha256 (both records), one added
  note on limit 11 (completed record) and the chain hashes changed.
- Addendum tavily-leads-vetting.json: the lead vetting of wf_6a6cb7b8-c22
  (Sonnet screen, then the sweep's facts refuter and two-family fit
  refuters). 136 leads in 30 layers, 100 dismissed, 36 proposed, 29 kept,
  0 survivors; the Claude fit refuter refuted all 29, GPT-6 21, the facts
  refuter 4; 8 split cases where GPT-6 did not refute but Claude did.
  Discovery-completeness evidence, not a verdict, with both search limits
  (Claude WebSearch capped: 2 calls, both capped; GPT-6 cached). Sanitized
  with host_receipts.sanitize(); one ref reads <checkout>. Its 23 40-hex ids
  are shortened to 12 characters, as in returns.json, because the file names
  sourcegraph/* leads (gitleaks sourcegraph-access-token rule). Its usage
  record child-usage-wf_6a6cb7b8-c22.json (usage_record.py): complete, 75
  children at effort max.
- README: GPT-6 cached web search and lead vetting sections, limit 11, files,
  redaction, verification; limit 11 and the addendum are not yet reviewed.

Verification, with a simulated registration of every new and changed file:
saturation_ledger.py --check --base origin/main passes (4 records, chain
intact) and scripts/validate.py passes. tests.test_saturation_ledger and
tests.test_landscape_sweep_harness: 141 run, OK, 2 skipped (no bash 3.2, no
shellcheck). tests.test_gitleaks_config and tests.test_verdict_lane_vendoring:
40 run, OK, once another scan's per-user gitleaks lock was free. Without the
registration, test_the_committed_ledger_checks fails only on unregistered
lane files, as before. manifests/evidence.json is not edited; the coordinator
registers the files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Register the landscape-sweep-20260926 record (hot-file commit)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Landscape-sweep tools: superseded attempts keep usage-integrity failures; failure coverage per worker and layer

Repairs the two P2 findings of a read-only GPT-6-Astra review (max effort,
live search) of the 2026-09-26 record's checkout. Each repair comes with a
regression test that failed before it.

(A) child-usage.mjs: a superseded attempt whose transcript held an assistant
message without provider usage left the run status "complete", although
by_resolved_model could not count that usage, so downstream usage could be
marked complete falsely. An attempt's usage-integrity failures (an assistant
message without provider usage or without a resolved model, or no transcript
file) are now its usage_issues; a superseded attempt with any makes the run
incomplete (the reason names them) and the CLI exit 1. Its other issues (no
result entry, the <synthetic> usage-limit row outside the family) stay
expected, and an attempt that made no request still counts as zero usage.
usage_record.py's summary names such attempts (superseded_usage_issues).
test-child-usage.mjs: four new checks failed before the repair (missing
usage, unresolved model, no transcript, CLI exit); a guard for an attempt
without requests passes. SHA256SUMS updated for the two changed files.

(B) make_result.py: failure coverage was checked over all layers at once, so
removing one layer's critic-cap failure and its reopen entry still passed
(another layer's "critic" failure satisfied the check) and that layer derived
a clean count of 1. effort_deviation and web_search_capped failures are now
checked per worker and per layer, with the mapping convert.py records them by
(deviation_rounds, moved unchanged to sweep_common.py): a
<role>:<layer>[:followup] worker's failure in that layer's round, the
critic's in every layer of the record; a worker naming no layer of the record
is refused as before. test_failure_coverage_is_checked_per_worker_and_per_layer
drops the critic's cap failure, then its effort deviation, from one of two
healthy layers, and moves a worker's failure to the other layer; all three
subtests failed before the repair.

Docs: the harness README (evidence contract, WebSearch cap, review repairs)
and the workflows README.

Tests: test-child-usage.mjs 51 passed, test-usage-receipts.mjs 18,
test-envelope.mjs 214, test-contract-mutations.mjs 47; test-codex-envelope.mjs
and check-syntax.mjs exit 0. tests.test_saturation_ledger,
tests.test_landscape_sweep_harness, tests.test_adoption_docs_consistency and
tests.test_verdict_lane_vendoring: 198 run, OK (3 skipped: no bash 3.2, no
shellcheck, no profile coverage change since the pinned release).
manifests/evidence.json is not edited; the coordinator registers the changed
files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Landscape sweep 2026-09-26 record: re-measured usage, GPT-6 review round, lane limit 12

Regenerated once with the pipeline of the earlier rounds, on the tools of
7760b1da, so the manifest sha changes once.

- Usage: the three usage records were re-measured with child-usage.mjs at
  7760b1da (sha256 85b95813...730d). Each child_usage is identical to the
  1bd2416b measurement: wf_8397ada1-777 stays complete (none of its 8
  superseded attempts holds uncounted usage) and run 1 stays incomplete, so
  lower_bound_usage and every ledger outcome are unchanged. Only
  tool_commit, tool_sha256 and measured_at_utc changed.
- Lane limit 12, a dated correction from the trading lane's review of #357:
  the execution-broker proposal of wboayue/rust-ibapi (a re-pin through
  nautilus_trader PR #5041) attributes #4983 to rc5. Checked with gh api:
  #4983's body reports Version v1.227.0; 1b0a49d2 is identical to tag
  v2.0.0rc5, and there crates/adapters/interactive_brokers/src/execution/
  core.rs:493-495 overrides handles_order_venue to return true, so the
  ClientVenueMismatch denial behind crates/execution/src/engine/mod.rs:2222
  cannot fire for the IB client; the open PR #280 reclassifies #4983 in
  catalogs/us-equities/runtime-target.json. Read instead: rc5 IBKR execution
  is unqualified (no rc5 stock order observed). The proposal text is kept.
- convert.py with limits 1-12: exit 3 on the same two notes, summary
  identical. returns.json and layers.json are byte-identical; lanes.json and
  manifest-20260926.json differ only by the added limit (pin ids unchanged:
  60 occurrences, 39 distinct).
- make_result.py with the per-worker, per-layer coverage check passes on
  this record and gives the same RESULT.json.
- Ledger: both 2026-09-26 records were re-appended onto origin/main's ledger
  (803bc35, unchanged since 3822178). Only manifest_sha256, usage_sha256
  and the chain hashes changed.
- README and independent-review.json: the third review round (GPT-6-Astra,
  NOT_READY, two P2 findings, both reproduced and repaired in 7760b1da) and
  the trading lane's correction.

Verification, with a simulated registration of the 18 changed registered
files: saturation_ledger.py --check --base origin/main passes (4 records,
chain intact) and scripts/validate.py passes. Tests: tests.test_saturation_ledger,
tests.test_landscape_sweep_harness, tests.test_adoption_docs_consistency and
tests.test_verdict_lane_vendoring 198 run, OK (3 skipped: no bash 3.2, no
shellcheck, no profile coverage change); test-child-usage.mjs 51 passed;
tests.test_gitleaks_config 27 OK. manifests/evidence.json is not edited; the
coordinator registers the changed files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Re-register the repaired sweep record (hot-file commit)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
…ily names in the leak checks

- codex_job.py passes -c web_search="live" instead of --search: --search before exec (and
  no flag) sends external_web_access false (cached index); only web_search="live" sends
  true (#332 artifacts websearch-*.json, W1-W3; a live call confirmed it on 2026-09-26).
- .claude/settings.json sets CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION=1500: the 2026-09-26
  sweep hit the default 200-per-session cap at 04:10Z, after which Claude workers got no
  WebSearch results (landscape-sweep-20260926 record, lane limit 10).
- blind-adjudicator.md (adoption and examples) and adjudication-prompt.md name astra,
  gpt-6-sol, gpt-6-luna, gpt-5.6-terra, fable and mythos; adjudicate.py's reported
  identity words add astra, fable and mythos (the gpt- form already covers the rest).
- The harness README documents both budgets. Tests: 312 OK.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
…ily names in the leak checks

- codex_job.py passes -c web_search="live" instead of --search: --search before exec (and
  no flag) sends external_web_access false (cached index); only web_search="live" sends
  true (#332 artifacts websearch-*.json, W1-W3; a live call confirmed it on 2026-09-26).
- .claude/settings.json sets CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION=1500: the 2026-09-26
  sweep hit the default 200-per-session cap at 04:10Z, after which Claude workers got no
  WebSearch results (landscape-sweep-20260926 record, lane limit 10).
- blind-adjudicator.md (adoption and examples) and adjudication-prompt.md name astra,
  gpt-6-sol, gpt-6-luna, gpt-5.6-terra, fable and mythos; adjudicate.py's reported
  identity words add astra, fable and mythos (the gpt- form already covers the rest).
- The harness README documents both budgets. Tests: 312 OK.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 26, 2026
…ude family names in the blind-adjudication leak checks (#352)

* GPT-6 live web search, WebSearch session budget, and GPT-6/Claude family names in the leak checks

- codex_job.py passes -c web_search="live" instead of --search: --search before exec (and
  no flag) sends external_web_access false (cached index); only web_search="live" sends
  true (#332 artifacts websearch-*.json, W1-W3; a live call confirmed it on 2026-09-26).
- .claude/settings.json sets CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION=1500: the 2026-09-26
  sweep hit the default 200-per-session cap at 04:10Z, after which Claude workers got no
  WebSearch results (landscape-sweep-20260926 record, lane limit 10).
- blind-adjudicator.md (adoption and examples) and adjudication-prompt.md name astra,
  gpt-6-sol, gpt-6-luna, gpt-5.6-terra, fable and mythos; adjudicate.py's reported
  identity words add astra, fable and mythos (the gpt- form already covers the rest).
- The harness README documents both budgets. Tests: 312 OK.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Mark the blind-adjudicator leak-check change after v2026.09.26.2 in bootstrap.md

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Register the new adjudication provenance for the leak-check change

adjudication_provenance() now hashes the changed adjudicate.py, adjudication-prompt.md and
blind-adjudicator.md; tests.test_verdict_lane_vendoring requires the current provenance in
tools/sota-convergence/lane-provenance.json (CI validate failure on #352).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Re-register after the provenance entry (hot-file commit)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:shared Touches files owned by both lanes; needs both lanes' acknowledgement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant