Repository navigation
Landscape-sweep GPT-6 lane through OmniRoute: lane-local Codex home with the token stack at max effort (stacked on #385) - #387
Merged
seathatflowsinourveins merged 10 commits intoSep 27, 2026
Conversation
…es never a refutation reason The sweep's templates change in two ways: - "Maintained upstream with 2025-2026 activity" becomes the OpenSSF Scorecard Maintained check. A stale repository (archived, or without a default-branch commit in the last 90 days) is excluded whatever its stars. A repository under 90 days old is flagged as too new to assess. An old tag on an active branch is not stale; pin the commit instead. - Licenses are recorded as information only and are never a reason to exclude, refute or rank down a candidate. This follows the operator's 2026-09-26 decision; the fit refuter was still told to refute non-commercial or unclear licenses. The template-hash change detector moves to PROMPTS_SHA256_CURRENT (dc5ffcb6...). The 2026-09-26 value stays as that run's recorded hash. A new test checks both rules are in the templates. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
…e the maintenance rule precisely Independent review (headless Sonnet 5, effort max): needs_changes, 1 major and 4 minor findings. - major: the shared common template, which reaches every role including the fit refuter, still judged "license and platform fit". It now reads "and platform fit (license is information only)". - minor: the rule is now "derived from" the Scorecard Maintained check and states Scorecard's actual criteria (archived is lowest; a 90-day window; top score needs a commit per week). This rule excludes only the zero-activity end. - minor: the commit fact is established with gh api .../commits?since=<90 days ago>, not pushed_at (any branch). - minor: a repository under 90 days old says "TOO NEW TO ASSESS (<90 days)" at the start of its demonstrated_gap. - minor: docs/harness-defaults.md's anti-pattern row no longer calls the template change pending. The regression test now scans every template for license-as-criterion phrases. PROMPTS_SHA256_CURRENT is updated to 0e443585. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
… necessary; REST URL alternative The independent Opus re-check agreed the major and 3 of 4 minor findings were resolved, and raised two minors. - The shared template said Scorecard "gives its top score only for at least one commit per week". It now gives the top score to at least one commit per week and partial credit for maintainers' issue activity (the Scorecard doc's sufficient condition). - The commit check also names the same https://api.github.com REST URL, for lanes without the gh CLI. The test comment says "derived from", and the anti-pattern log names the test. PROMPTS_SHA256_CURRENT is updated to a9722fec. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
…y lanes cannot reach the commit evidence) GPT-6 cross-family review (gpt-6-astra, max) of b87a656 returned needs_changes, with 1 medium finding. The web-only GPT-6 lane cannot run gh (make_prompt.py TAIL), and its web tools failed to fetch the commits API URL. Since the fit refuter defaults to refuted, maintained repositories could be refuted for unreachable evidence. - common: a lane that cannot reach either source records the commit fact as unknown and never excludes or refutes on unknown maintenance, because a missing result is pending, never refuted. - fit: staleness must be established from evidence; unknown maintenance is never a reason to refute. The test asserts both. PROMPTS_SHA256_CURRENT is updated to f64eec22. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
… with the token stack at max effort The operator's 2026-09-27 direction: route the sweep's GPT-6 lane through OmniRoute, with the full token-efficiency practice, then run the 20-layer sweep through it. build_args.py --gpt6-provider omniroute --codex-host <HOST> stages <work-dir>/codex-home: - config.toml: an OmniRoute provider block (cx/gpt-6-astra, effort max, loopback /v1, env_key OMNIROUTE_API_KEY, wire_api responses, no websockets), plus the [mcp_servers.*] tables of codex.config.template.toml rendered for the host by the checkout's render_config.py (serena, ai-memory, socraticode, headroom, codebase-memory, qmd, context-mode). Project and hook trust are left out. - stack-worker.config.toml: the Codex worker profile, copied verbatim. codex_job.py runs that lane with CODEX_HOME set to the staged home, without --ignore-user-config, and with -p stack-worker. The key comes from the OMNIROUTE_API_KEY variable. A keyless loopback gateway gets the "local-loopback" placeholder, and a real key wins. The quota gate is refused for a gateway lane, and inputs.json records the provider. Model names may carry a provider prefix. The native default is unchanged. Tests: 8 new, covering rendering, the runner's argv and env, key handling and refusals. 161 run, OK (3 skipped). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
Real hosts' value files are gitignored (adoption/hosts/*.json), so they exist only in the checkout that owns them, never in a worktree. --codex-host now takes either a host name (adoption/hosts/<name>.json, through render_config's load_host_values) or a path to a *.json host value file, checked for the same flat string shape. The lane record and the config.toml header carry only the file's stem, never its private location. Staged for this workstation with its private file, the lane home is provider omniroute, cx/gpt-6-astra, effort max, profile stack-worker and the seven token MCP servers, and it is valid TOML. Tests: one new; 162 run, OK (3 skipped). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
…ail-closed MCP extraction The GPT-6 cross-family review of 5b43474 returned needs_changes: 1 high, 2 medium and 1 low finding. - high: Codex's default shell snapshot wrote the exported environment, the provider key included, into codex-home/shell_snapshots/*.sh at mode 0644 (shell_snapshot_exports.rs at rust-v0.157.1). The lane config now sets [features] shell_snapshot = false. - medium: GPT-6 Astra runs Responses Lite, which carries no hosted tools, and custom providers default to no standalone search. So the lane gets no web search. The provider now advertises supports_standalone_web_search = true (Codex advanced config docs) and [features] standalone_web_search = true (under development in 0.157.1). The parity check must still show that OmniRoute serves the endpoint. - medium: the default worker profile arrives with the Codex worker-lane change, so the error now names --stack-worker-profile. The README states the dependency. - low: the MCP extraction accepts quoted and spaced table headers, and is checked against a real TOML parse. An extraction that parses badly or names different servers than the whole rendered config fails closed. A blank key value counts as unset, since Codex rejects a blank env_key. The README records the isolation limits: a trusted cwd .codex/config.toml and the system config still layer in; the runner's cwd is empty, and this workstation has no /etc/codex/config.toml. Tests: 164 run, OK (3 skipped). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
Base automatically changed from
claude/sweep-templates-maintained-license-20260927
to
main
September 27, 2026 06:40
…e-omniroute-20260927 # Conflicts: # manifests/evidence.json
The GPT-6 re-check of 8780d93 resolved the snapshot and search findings. It left one low finding open: after [mcp_servers.demo], an env table written as [mcp_servers.'demo'.env] or [ mcp_servers . demo . env ] was dropped silently, because the check compared server names only. - The check now compares the parsed mcp_servers table of the extraction with the whole rendered config's, nested settings included. Any difference names the servers and fails staging closed. - The header pattern accepts TOML's literal-quoted names and spaces around the dots, so those spellings are kept rather than refused. - Without tomllib (Python 3.11+), staging refuses instead of skipping the check. The failing control: both new tests fail against 8780d93's build_args.py. The real template rendered for this host still extracts all 7 servers, env tables included, equal to the full parse. The other open medium (default worker profile absent) resolves when #389 lands the profile, and #389 merges first. Tests: 166 run, OK (3 skipped); validate.py passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
Owner
Author
|
GPT-6 cross-family re-check of the repair (8780d93): needs_changes. Findings and their state at 1e81cce:
What 1e81cce changes:
Also from the re-check: no new independent defects. An absolute Checks at 1e81cce: Merge order stays #389 → this PR → parity probe on the fixed OmniRoute build → full 20-layer foundation sweep. 🤖 Generated with Claude Code |
…e-omniroute-20260927
seathatflowsinourveins
deleted the
claude/sweep-gpt6-lane-omniroute-20260927
branch
September 27, 2026 07:36
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 27, 2026
… are an allowed dispatch mode (#388) Practice gap #6: tools/adoption/install_skills.py now says what it checked. The SKILL.md sha256 matches the pin, and the lock records the tree. Practice gap #8: agent teams are an allowed dispatch mode, per the current Claude Code docs (agent-teams, context and communication). - examples/claude-native/CLAUDE.md gains one word-neutral sentence (924 words). - scripts/adoption_status.py reports agent_teams_opt_in (0 or 1) as information, not agent_teams_off as a wiring requirement. - docs/token-efficiency-stack.md is updated to match. - docs/harness-rules-convergence-20260922.md marks WH-3 superseded. GPT-6 cross-family review: accept, no findings (213 tests, OK). After main (#389, #387) was merged in: validate.py passed; tests.test_adoption_status and tests.test_landscape_sweep_harness OK; CI full suite green. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
This was referenced Sep 27, 2026
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 27, 2026
…count pool catalogs/foundation/manifest.json, native-clients selection: records the user's 2026-09-27 decision. A loopback OmniRoute gateway is adopted for pooled GPT-6 access by gateway-routed Codex lanes such as the landscape sweep. Native Codex stays the max-quality default. The agent-sdks layer default stays pending the preregistered three-arm workers comparison (docs/grand-catalog-handbook.md:602), which remains its only acceptance gate. Model routing is still not a separate layer. Only layer prose changes; checked_at and the decisions file are untouched. adoption/credential-inventory.json, omniroute entry: - a keyless loopback gateway with an optional per-lane key; - callers pass the placeholder local-loopback, and the store stays absent; - a host or lane that turns REQUIRE_API_KEY=true creates scoped keys per lane; - tools/sota-convergence/landscape-sweep/codex_job.py (#387) and Codex's env_key are added as environment-only consumers; - the lane is now model-gateway, because the foundation's Codex lanes consume the key as well as the trading routing recipe. The variable set is unchanged, so the secret-path guard's name list still covers every inventory variable. The docs/secret-storage.md row matches the entry. The decision record gains one line: the template adds an explicit OMNIROUTE_SERVER_HOST=127.0.0.1. Sources: - the decision record; - https://github.com/diegosouzapw/OmniRoute at a58000c7: bin/cli/utils/serverHost.mjs L16-26, docs/guides/CODEX-CLI-CONFIGURATION.md ("Local unauthenticated OmniRoute"), src/shared/validation/schemas/keys.ts (key scoping fields). Checks: validate_foundation.py errors [] and 159 tests OK across test_foundation_catalog, test_secret_path_guard, test_host_requests, test_ecosystem_manifest and test_credential_status. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 27, 2026
…plate and stack records for the release/v3.8.51 source build (#396) * OmniRoute account pool: decision record and sanitized evidence Records the gateway the coordinator installed on the workstation on 2026-09-27 at the user's direction. The build: - OmniRoute release/v3.8.51 at a58000c7 plus upstream PR #14904 (the Proxy-safe deadline wrapper) and PR #13788 (/v1/alpha/search for Codex web.run); - built with upstream's own scripts, BUILD_SHA dd6e9607e. The installation: - a systemd --user service on loopback; - a keyless, passwordless posture, with its risk recorded; - Codex wired through an omniroute provider and profile at max effort. The record says plainly that the preregistered three-arm workers comparison stays the gate for any agent-sdks default change. Evidence classes stay separate: - live provider execution: probe run 2, the gateway's effort columns, and the search probe with its log; - unchanged upstream tests: 26 route and keepalive files, 184 of 186 pass, 2 skipped; - local integration: the settings read-back, the installed unit, the build identity, and cherry-pick patch-ids with the build tree; - a synthetic Proxy reproduction on three Node runtimes; - live upstream metadata: PR, issue, npm and CI state. Reported but not retained: the red-to-green keepalive run, the alpha-search test and the peer's 8-check probe. The full upstream test:unit summary is pending. Sources: - https://github.com/diegosouzapw/OmniRoute at a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3; - pulls 14904 and 13788, issue 14866; - the files cited per decision. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * foundation-stack: scope the 3.8.50-only gateway claims; add the v3.8.51 source build Two sentences in the OmniRoute section held only for 3.8.50, and are now scoped to it: - the executor's xhigh clamp for unknown models (open-sse/executors/codex.ts L346-347 at 5458026c); - "native Codex passthrough skips router compression", a bypass that a58000c7 removed (open-sse/handlers/chatCore.ts L1429-1449 at a58000c7, whose comment names the codex/* exclusion as the remedy). A short section describes the workstation's source build of release/v3.8.51 at a58000c7 with #14904 and #13788. It links the decision record, the evidence, upstream's build scripts and the new unit template. It also covers effort, compression, the cx/ slug with no router alias, and the keyless loopback posture. The component pin stays 3.8.50, so the page's release links still match manifests/stack.json (tests/test_foundation_stack_pins.py). Sources: - https://github.com/diegosouzapw/OmniRoute at a58000c7 and 5458026c; - pulls 14904 and 13788. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * OmniRoute gateway: values-free systemd --user unit template with structural tests adoption/templates/systemd/omniroute.service mirrors the unit installed on the workstation (the evidence copy). Three placeholders stand in for host values, following this directory's @name@ convention: @OMNIROUTE_PREFIX@, @NODE_PREFIX@ and @CODEX_CLIENT_VERSION@. Secrets arrive only through EnvironmentFile=. The file holds plain NAME=value lines, the format upstream's own scripts/build/bootstrap-env.mjs writes to {DATA_DIR}/server.env. Every non-secret Environment= line carries its reason and upstream source. The template adds one line the installed unit lacks: Environment=OMNIROUTE_SERVER_HOST=127.0.0.1. `omniroute serve` binds 0.0.0.0 when that variable is unset (bin/cli/utils/serverHost.mjs L16-26 at a58000c7), and the keyless posture needs loopback. On the workstation the environment file already sets the same value. tests/test_omniroute_gateway_unit.py checks, offline: - the template, rendered with the workstation's values, reproduces the recorded installed unit; - the placeholder set is exact and documented; - secret-like names never appear in Environment=; - each Environment= line has its own reason; - serve runs in the foreground (no --daemon or --no-recovery) with Restart=on-failure. Discriminating controls plant a secret line, an unexplained line and two drifts, and each is caught. Manual check, rendered copy: systemd-analyze --user verify rc 0 (systemd 255). A copy with a missing ExecStart binary fails with rc 1. Sources: - https://github.com/diegosouzapw/OmniRoute at a58000c7: bin/cli/commands/serve.mjs, bin/cli/utils/serverHost.mjs, bin/cli/utils/pid.mjs, scripts/build/runtime-env.mjs, scripts/build/bootstrap-env.mjs, src/shared/utils/runtimeTimeouts.ts, src/shared/services/cliRuntime.ts; - systemd.exec(5), EnvironmentFile=. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Foundation catalog and credential inventory: the adopted OmniRoute account pool catalogs/foundation/manifest.json, native-clients selection: records the user's 2026-09-27 decision. A loopback OmniRoute gateway is adopted for pooled GPT-6 access by gateway-routed Codex lanes such as the landscape sweep. Native Codex stays the max-quality default. The agent-sdks layer default stays pending the preregistered three-arm workers comparison (docs/grand-catalog-handbook.md:602), which remains its only acceptance gate. Model routing is still not a separate layer. Only layer prose changes; checked_at and the decisions file are untouched. adoption/credential-inventory.json, omniroute entry: - a keyless loopback gateway with an optional per-lane key; - callers pass the placeholder local-loopback, and the store stays absent; - a host or lane that turns REQUIRE_API_KEY=true creates scoped keys per lane; - tools/sota-convergence/landscape-sweep/codex_job.py (#387) and Codex's env_key are added as environment-only consumers; - the lane is now model-gateway, because the foundation's Codex lanes consume the key as well as the trading routing recipe. The variable set is unchanged, so the secret-path guard's name list still covers every inventory variable. The docs/secret-storage.md row matches the entry. The decision record gains one line: the template adds an explicit OMNIROUTE_SERVER_HOST=127.0.0.1. Sources: - the decision record; - https://github.com/diegosouzapw/OmniRoute at a58000c7: bin/cli/utils/serverHost.mjs L16-26, docs/guides/CODEX-CLI-CONFIGURATION.md ("Local unauthenticated OmniRoute"), src/shared/validation/schemas/keys.ts (key scoping fields). Checks: validate_foundation.py errors [] and 159 tests OK across test_foundation_catalog, test_secret_path_guard, test_host_requests, test_ecosystem_manifest and test_credential_status. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * OmniRoute evidence: the upstream test:unit result on the dd6e9607e tree The coordinator's upstream `npm run test:unit` run finished at 08:17Z (Node 24.21.0, upstream's own runner flags). Its summary is added verbatim, together with an extraction by the record's author: - the runner's own totals: 43,522 tests, 43,460 pass, 31 fail, 31 skipped; - the 21 failing files and the 31 failing names; - a diff with the 30 names from the same files on bare a58000c7. Exactly one failure comes from the picks. #13788's /v1/alpha/search route adds a connection-query site that tests/unit/hard-session-lease-bypass-inventory.test.ts:348 does not classify. It is not patched locally. The other 30 are base reds of the release branch (#14866). #14904 adds none. The output holds one totals block. upstream's test:unit chains three stages with &&, so stage 1's failures stopped the chain: the dashboard stage and test:unit:serial did not run (package.json line 132 at dd6e9607e). The README and the decision record say so, and do not call this the full suite. Sources: - https://github.com/diegosouzapw/OmniRoute package.json "test:unit" and "test:unit:serial" at the build tree; - the tests named above. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * OmniRoute record: fixes from the GPT-6 cross-family review and the full unit run One GPT-6 review round (codex exec, read-only) reported 5 medium and 2 low findings. All 7 are fixed here. 1. medium: gateway-search-log.jsonl was ignored by the repository's *.jsonl rule, so it was never committed or registered. Fix: a narrow .gitignore exception for this evidence directory, following the existing per-directory exceptions, and the file is now tracked. 2. medium: proxy-repro.txt shared one POST body between its two arms. Its Node 22 "already been used" result came from that shared body, not from the Proxy. Fix: a corrected rerun, proxy-repro-independent. txt, gives each arm its own Request. Proxied POST and GET fail with the #state TypeError on Node 24.21.0 and 26.10.0, and construct on Node 22.23.3. The first attempt is kept and its flaw noted. The decision now says defect 1 affects the Node 24 and 26 lines, that upstream's Docker base is node:26 (Dockerfile L2), and that a Node 22 gateway was not tried. #13788 needed a source build anyway. 3. medium: the claim that secrets arrive only through EnvironmentFile= overstated what a unit can enforce, because a user service inherits the user manager's environment (systemd.exec(5)). Fix: the template, decision, foundation-stack and README now say the unit holds no secret and adds its secrets only through EnvironmentFile=. They also say how to keep credentials out of the manager. The test is renamed to what it checks. 4. medium: the route and keepalive summary lacked its selection and skips. Fix: upstream-tests-route-keepalive-detail.txt retains the totals, the two upstream skips with their reason, 0 not-ok lines, and #14904's two passing cases. The command line and the 26-file list are labelled reported, not retained. 5. medium: the decision cited research rows and an X1 preregistration that are not in the repository. Fix: every such claim now carries a pinned upstream citation, checked against local copies. That covers: - the divergences: codex.ts L1467-1484, L1362-1369, L366, L374-393; Codex client.rs L859-878, L897-954; provider.rs L410-423; - the risk: management.ts L261-266, routeGuard.ts L79, hooks route.ts L79; - the backups: backup.mjs L27-32; - systemd's env-util.c L28-50 at v255. The synthesis items are labelled as recommendations not retained here, and X1 is described as a sketch that is not frozen and has not run. 6. low: settings.ts forces requireLogin only while first-time setup is incomplete. Fix: the decision now states the condition. The verbatim make_passwordless.py record is annotated in the README rather than edited. 7. low: the suite summary's 07:45-08:17Z range was inconsistent. Fix: a dated correction from the process's elapsed time (14:43 at 07:45:47Z) and the output file's mtime (08:03:29Z) puts the run at about 07:31-08:03Z, which matches the runner's 1944 s. The coordinator's summary stays verbatim. Also: - the decision's scope line no longer lists an observation-inference change that was not made; - one nested list is re-indented; - the EnvironmentFile= exposure paragraph cites bin/omniroute.mjs L134-216 and serve.mjs L281-298, and states the departure from docs/secret-storage.md step 9. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * OmniRoute record: precision pass on the template's reasons and three citations A line-by-line re-check of every cited upstream line, after the GPT-6 review round, found one misquote and three claims without a citation for their mechanism. GPT-6 did not see this commit (one review round only). Each change below was read at its pin: the local a58000c7 checkout, and Codex rust-v0.157.1 and systemd v255 through the GitHub contents API. Decision record: - Sticky limit. The record said AUTO-COMBO.md recommends sticky 1 for "one-model rotation". Upstream says to set the combo override to 1 "for one-request rotation" (docs/routing/AUTO-COMBO.md L354-357). The codex provider strategy's limit is read before the global setting (src/sse/services/auth.ts L1999-2000). - export lines. systemd's parser keeps "export NAME" as the key (src/basic/env-file.c L75-95). The EnvironmentFile= loader then drops every assignment with an invalid name (src/core/execute.c L773-787; src/basic/env-util.c L542-554, L78-90, L28-50). The record cited only the name check before. - Stream timeouts. The readiness budget never exceeds its cap (open-sse/utils/streamReadinessPolicy.ts L198). The content-stall watchdog reuses that budget (open-sse/handlers/chatCore.ts L6395-6399). The active timeout is a hard cap that upstream bytes never reset (open-sse/config/constants.ts L31-33). - The template mirrors the installed unit's directives apart from Description=. The source-review line names every pin. Unit template and its test: - Each Environment= line now has a one-line reason, as the unit brief asks. The lsof shim's failure and removal text moves to the header. The header says Description= and comments also differ from the installed unit. - unexplained_environment() now requires exactly one reason line. A planted two-line reason is caught. The previous template fails the stricter check on 3 lines: PATH, OMNIROUTE_SERVER_HOST and STREAM_READINESS_TIMEOUT_MS. Directives are unchanged, so the template still renders to the recorded installed unit. Evidence README: the installed-unit row no longer says every Environment= line has a reason comment. OMNIROUTE_MEMORY_MB has none in the installed unit. Sources: - https://github.com/diegosouzapw/OmniRoute at a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3; - https://github.com/systemd/systemd at v255; - the repository's own tests/test_token_report_refresh_units.py style for planted-violation controls. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * OmniRoute evidence: scripted probe records, build per observation, listener and search-row observation The independent verifier found that probe-lane-run2.json and probe-search-profile-only.json were derived by hand in an undeclared step, and that this step dropped one of run 2's four gateway rows. Changes: - scripts/derive_probe_records.py.txt re-derives both files from the same private outputs. The run 2 record keeps all four rows, with conn_hash replaced by an account letter, plus the rollout turn_context line and effort_fields as printed. The script also writes probe-lane-client-events.json: Codex's side of the 06:32Z "high demand" probe and of run 1. The README declares the sanitization, as docs/acceptance-evidence-policy.md requires. - The README explains the four rows. Each request is listed twice, once as an in-memory entry and once as a persisted row. That matches OmniRoute's list route (src/app/api/usage/call-logs/route.ts L116-226 and src/lib/usage/callLogs.ts L457-508 at a58000c7). - build-delta.txt shows that bf0255649 and dd6e9607e differ in two files. The README now names the build behind every live observation. - gateway-readonly-observation.txt comes from scripts/observe_gateway.py.txt, a read-only run at 10:40Z. The unit's only listeners are 127.0.0.1:20128, :20131 and :20132, and call_logs holds one /v1/search row per SEARCH line. That corrects the earlier claim that /v1/alpha/search writes no call_logs rows (open-sse/handlers/search.ts L1625-1638). - Two more items are reported, not retained: route_repro's status lines and run 1's account relation. Run 1's follow-up has no call_logs row, and the README says so. Sources: - https://github.com/diegosouzapw/OmniRoute at a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3 for the files above, and dd6e9607e4884ec75c9bc0d96e60b01e9483d84e for the build delta; - https://github.com/openai/codex at rust-v0.157.1, codex-rs/codex-api/src/api_bridge.rs L157-158 and codex-rs/protocol/src/error.rs L160-161 (HTTP 500 is shown as "high demand"); - iproute2 ss(8), https://man7.org/linux/man-pages/man8/ss.8.html, and the cgroup v2 cgroup.procs file, https://docs.kernel.org/admin-guide/cgroup-v2.html, for the listener observation; - docs/acceptance-evidence-policy.md ("declare every public sanitization"). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * OmniRoute record: verifier fixes to the decision, the recipe, the unit template and its test Decision record: - It names the build behind each live observation. The lane probe and the effort rows ran on bf0255649. Their results carry to dd6e9607e by the two-file build delta, not by a rerun. - The loopback bind is now cited for all three listeners, and the 10:40Z observation backs it. LIVE_WS_HOST, EMBED_WS_PROXY_HOST and the two _PORT variables join the keep-out-of-the-manager guidance. - It gives the source for Codex's "high demand" message, and lists route_repro's status lines and run 1's account relation as reported, not retained. - "No API exposes these columns" now rests on mapSummaryRow, and L646-653 explain when the columns are filled. - The CODEX_CLIENT_VERSION citation now points at open-sse/config/codexClient.ts, and the facts the verifier found uncited now carry their lines. The "every factual claim" sentence is narrowed. - The claim that /v1/alpha/search writes no call_logs rows is corrected. docs/foundation-stack.md: the Node 26 failure is labelled as a synthetic-fixture result, and the effort observation names its build. Unit template: the manager guidance covers the WebSocket bind variables, and the CODEX_CLIENT_VERSION reason cites the right file. The directives are unchanged, so the template still mirrors the installed unit. tests/test_omniroute_gateway_unit.py: the EnvironmentFile= count, the Environment= name set, the credential-text regex, the placeholder set and the supervision settings are now helpers with planted controls. A recorded red run fails the main secret test on a second EnvironmentFile=, an undocumented Environment= name and an export line, and passes it on the clean template. Sources, OmniRoute at a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3 (https://github.com/diegosouzapw/OmniRoute): - src/server/ws/liveServer.ts L50-54, L713-714; - src/lib/services/embedWsProxy.ts L35-36, L240-242, L256-257; - .env.example L2302-2308; - open-sse/config/codexClient.ts L13, L29-34, L36-83; - src/shared/constants/codexClient.ts L6; - src/lib/usage/callLogs.ts L457-508, L646-653, L1021, L1040; - src/app/api/usage/call-logs/route.ts L116-226; - open-sse/handlers/search.ts L1625-1638; - bin/cli/locales/en.json L255-256; - bin/cli/commands/serve.mjs L64-65; - src/sse/services/auth.ts L1995-2001; - src/lib/db/settings.ts L159, L163; - open-sse/services/compression/types.ts L421-423; - src/lib/db/compression.ts L663-664; - open-sse/services/thinkingBudget.ts L8-9, L19, L65-67. Also https://github.com/openai/codex at rust-v0.157.1: codex-rs/codex-api/src/api_bridge.rs L157-158 and codex-rs/protocol/src/error.rs L160-161. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * OmniRoute record: precision on the listing pairing, the effort-column APIs and the manager bind variables - The four-row listing pattern is stated as consistent with the list route's merge, not proven, because the ids are withheld. - The in-memory entries of the list route also carry no effort field (src/app/api/usage/call-logs/route.ts L141-174, L186-215 at a58000c7). - A LIVE_WS_HOST or EMBED_WS_PROXY_HOST in the user manager could move a listener off loopback; it does not always do so. Source: https://github.com/diegosouzapw/OmniRoute at a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * OmniRoute unit test: the template-only and render-away checks become helpers with planted controls The module docstring says every check of the template's text is a helper that a planted violation fails. Two main assertions were still direct: the TEMPLATE_ONLY presence and absence loop, and the render-away check. Now: - template_only_problems() replaces the loop. Its control drops the line from the template and adds it to the installed copy. - placeholder_problems() also reports placeholders left after rendering. The unknown-placeholder control now expects both findings. A recorded red run fails each main template test on a planted copy, and each passes on the clean template (units/gateway/round2/red-run.txt). Source: docs/acceptance-evidence-policy.md, "Discriminating controls". Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * OmniRoute record: #13788 is an upstream feature, not a fix The build carries one upstream fix (#14904, the deadline wrapper) and one upstream feature (#13788, the /v1/alpha/search route for Codex's standalone web search). Reword "two upstream fixes" and "both fixes" in the decision title, the foundation-stack section, the foundation catalog status and the upstream test summary. Raised by the other lane's acknowledgement review of #396. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Shared hot files: omniroute component records the running build and the workstation posture; evidence registration manifests/stack.json, omniroute. The schema has no field for a source build with cherry-picks, and scripts/landscape.py requires version and source_pin to equal the dated landscape freshness snapshot. So the pin stays the published 3.8.50 (5458026c, still npm latest on 2026-09-27). freshness records the running build: - the base: release/v3.8.51 at a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3; - the picks: #14904 and #13788, with patch-ids equal to upstream's; - BUILD_SHA dd6e9607e; - a pointer to the decision record and the evidence directory. Other fields: - command_scope names the unit template and the adoption scope. It now says that the workstation gateway runs keyless and passwordless on loopback by the user's decision (decision 5), and that a multi-user or non-loopback host keeps the recipe's API key and dashboard login. Native Codex stays the max-quality default, and agent-sdks waits on the three-arm comparison. - upstream_sources gains the base commit and the two PRs. - role is unchanged. tests/test_stack_lifecycle.py L31-32 pins it to the row in blueprints/token-native-focus/saturation-audit.json, so the stale "authenticated" wording is a recorded residual. manifests/evidence.json, per the docs/lanes.md hot-file protocol: - taken from main at 5f3a7c2; - re-registers the changed tracked files: .gitignore, the credential inventory, the foundation catalog manifest, foundation-stack, secret-storage and stack.json; - registers the 35 evidence files, the decision record, the unit template and its test. That is 44 files, the full list of files this branch adds or changes; - component_matrix.py --write and new_host_grand_list.py --write were rerun, with no content change. validate.py passed with hashed_files 7288. Sources: - https://github.com/diegosouzapw/OmniRoute at commit a58000c7 and pulls 14904 and 13788; - docs/lanes.md "Hot-file protocol". Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
4 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
This PR routes the landscape-sweep harness's GPT-6 lane through the local OmniRoute gateway, at max effort and with the full token stack. Operator direction, 2026-09-27: "run the full 20 layer sweep, full parallel but first resolute with the omniroute wire with full token effiency practice first and we proceed via them".
build_args.py --gpt6-provider omniroute --codex-host <HOST>stages a lane-local Codex home,<work-dir>/codex-home/, with two files:config.toml: the OmniRoute provider block, plus the token-stack[mcp_servers.*]tables ofadoption/templates/codex.config.template.toml, rendered for the host by the checkout's owntools/adoption/render_config.py(load_host_values,render_one). The servers are serena, ai-memory, socraticode, headroom, codebase-memory, qmd and context-mode.model = "cx/gpt-6-astra",model_provider = "omniroute",model_reasoning_effort = "max", and[model_providers.omniroute]with a loopbackbase_urlending in/v1,env_key = "OMNIROUTE_API_KEY",requires_openai_auth = falseandwire_api = "responses".supports_websocketsstays unset: OmniRoute's WebSocket bridge drops the client-version headers that its HTTP/v1/responsespath forwards.stack-worker.config.toml: the Codex worker profile, copied verbatim, from--stack-worker-profileoradoption/templates/codex.stack-worker.config.toml.codex_job.pyruns a gateway lane withCODEX_HOME=<work-dir>/codex-home, without--ignore-user-config(which would drop that home'sconfig.toml), and with-p stack-worker.$OMNIROUTE_API_KEYin the harness's environment.omniroute setup --non-interactive,REQUIRE_API_KEY=false), the staged placeholderlocal-loopbackfills an unset variable, because Codex'senv_keyonly needs the variable to exist. A real key in the environment wins.--omniroute-require-keystages no placeholder; a job without the variable then ends with exit 6 before codex starts.--quota-stop-percentis refused for a gateway lane, since the probe reads the native login.--omniroute-base-urlmust be loopback.inputs.jsonrecords the provider, so a gateway run never reuses a native result.Parity against the live gateway, 2026-09-27 07:02Z. The gateway build was OmniRoute 3.8.51 = release/v3.8.51
a58000c7+ upstream PR #14904 (BUILD_SHAbf0255649). The lane was staged from this branch merged with #389, using the committed default profile.codex execexit 0 in 29 s;ctx_execute(context-mode), which returnedmcp-ok;--output-schemaoutput;turn_contexteffortmax;call_logs, whose reasoning row readsreasoning_effort_requested=max,reasoning_effort_upstream=max.web.runPOSTs<base_url>/alpha/search, and this build answers404 Unknown API route: /v1/alpha/search. The upstream implementation is diegosouzapw/OmniRoute PR #13788: open,deferred-v3.8.52, base release/v3.8.51, 2 new files. It uses OmniRoute's search registry, with keyless DuckDuckGo-lite as the fallback, not OpenAI's hosted search. The gateway owner is cherry-picking it. The check is re-run before the full sweep, and the sweep records the search backend as a lane limit.Review
GPT-6 cross-family review (
gpt-6-astra, effort max, read-only), two rounds:shell_snapshots/*.shat mode 0644. Nowshell_snapshot = false.supports_standalone_web_searchandstandalone_web_search.Tests
OmniRouteLaneBuildTests: a real render of the template foradoption/hosts/example.json; all 7 MCP servers, the provider block and valid TOML (tomllib); no trust tables and no unrendered${…}; the refusals; the require-key option; the native default unchanged.OmniRouteLaneRunnerTests: the exact argv; the environment'sCODEX_HOME; the key present; no start without a key; the placeholder used only when the variable is unset.python3 -m unittest tests.test_landscape_sweep_harness tests.test_saturation_ledger: 166 run, OK (3 skipped). Withtests.test_codex_worker_laneafter merging main (PR-D: Codex worker lane: Codex's own config writer, RTK text inline with exceptions, max-effort stack-worker profile #389): 206 run, OK (5 skipped).scripts/validate.pyandscripts/build_ecosystem.py --check: passed.SOTA sources
model_providers,env_key,wire_api,requires_openai_auth,CODEX_HOME, profiles): https://developers.openai.com/codex/config-reference$CODEX_HOME/<name>.config.tomlfor--profile <name>:codex-rs/config/src/config_layer_source.rsat rust-v0.157.1, as cited by the stack-worker profile.a58000c7, as verified by the token lane: https://github.com/diegosouzapw/OmniRoutetools/adoption/render_config.py.🤖 Generated with Claude Code
https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5