Skip to content

GPT-6 gateway lane on the all-engines OmniRoute instance: static headers and a one-slash model - #439

Merged
seathatflowsinourveins merged 2 commits into
mainfrom
claude/gpt6-lane-fw-instance-20260927
Sep 27, 2026
Merged

seathatflowsinourveins merged 2 commits into
mainfrom
claude/gpt6-lane-fw-instance-20260927

Conversation

@seathatflowsinourveins

@seathatflowsinourveins seathatflowsinourveins commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Scope

  • What this PR changes: the landscape-sweep GPT-6 lane can target the framework OmniRoute instance on 127.0.0.1:20129.
    • That instance has all engines and output styles on and is chained through node sharedgw to the shared gateway on 20128.
    • The shared 20128 stays unchanged for the memory services and the other sessions.
  • Base commit: 7fa3849c (Pre-push registry gate: run the three registry tests on every pushed commit #438). Reviewed as 55fc8d17..854803b0; the rebase changed no reviewed line, and docs/harness-defaults.md keeps Pre-push registry gate: run the three registry tests on every pushed commit #438's row beside this PR's two rows.
  • Head: 74c233d7
  • Lane: lane:foundation
  • Paths (6 files):
    • tools/sota-convergence/landscape-sweep/{build_args.py,codex_job.py,README.md}
    • tests/test_landscape_sweep_harness.py
    • docs/harness-defaults.md (two anti-pattern rows)
    • manifests/evidence.json (registration)
    • No generated report changes.

Changes

  • Static provider headers. build_args.py --omniroute-header NAME=VALUE stages Codex's model_providers.<id>.http_headers, for example x-omniroute-compression=allow-lossy.
    • Only OmniRoute 3.8.51's four per-request switches are accepted: x-omniroute-compression, -no-cache, -no-memory and -strip-reasoning.
    • Other x-omniroute-* headers can carry a secret under a name with no credential word, so a word filter cannot keep them out: x-omniroute-self-hop, -video-bridge-broker and -lease-owner, for example. Staged values are recorded in staged.json and in each job's inputs.json.
    • A refused name or value is never echoed.
    • codex_job.py applies the same allowlist to staged.json. It checks the lane home's table against it and records the headers per job.
    • The runner refuses staged headers when tomllib is missing (Python 3.9 and 3.10). A header-less lane still runs on 3.9. The builder's read-back is unconditional, because staging already needs 3.11.
  • The model has at most one provider segment, of letters, digits, _ and -. On 20129 it is sharedgw/gpt-6-astra-max.
    • Codex rust-v0.157.1 strips one namespace, and only one of those characters, when it looks up model metadata (codex-rs/models-manager/src/manager.rs L763-780).
    • Any other slug gets generic fallback metadata (model_info.rs L99-150): a different prompt template, no Responses Lite, no multi-agent, a 272k context window and a 10,000-byte tool-output truncation. That covers sharedgw/cx/gpt-6-astra and my.gw/gpt-6-astra-max.
    • main already refused a second slash. This PR adds the fallback-metadata diagnostic and the namespace character rule to both entry points.
  • README:
    • Manual commands must pass -m sharedgw/gpt-6-astra-max, because under -p stack-worker the profile's model otherwise wins.
    • The parity check compares the rendered items' content, without id and create_time and with the model masked. Its commands were run as written (evidence below).
    • Slashless aliases: the 401 observation is kept, without a mechanism, since none was established from source.
    • Effort evidence:
      • A gateway's outbound request is pipelinePayloads.providerRequest in its detailed call log, when detailed pipeline logging is on.
      • requestBody is the body that gateway received.
      • A non-null reasoning_effort_upstream is positive evidence; only a null effort column proves nothing.
  • Anti-pattern rows:
    • The two-slash slug on a chained gateway, checked through both real entry points.
    • A null call_logs effort column read as "effort not sent", and the first correction's inbound field. No repository check enforces this; the row says so.

SOTA sources

  • openai/codex rust-v0.157.1:
    • codex-rs/models-manager/src/manager.rs L763-780: the namespace rule. model_info.rs L99-150: the fallback metadata.
    • codex-rs/core/src/prompt_debug.rs: prompt-input.
    • codex-rs/model-provider-info/src/lib.rs:
      • L166-168: http_headers, whose values are a RedactedString.
      • L385-412: build_header_map skips an invalid name or value.
      • L135: deny_unknown_fields is a JSON-schema attribute only.
  • OmniRoute 3.8.51 (build dd6e9607e), installed source:
    • The switches:
      • open-sse/handlers/chatCore/headers.ts: L18-33 no-memory, L35-46 compression, L48-65 strip-reasoning.
      • src/lib/semanticCache.ts L483-486 and L501-504: no-cache.
      • open-sse/services/compression/lossyRequestPolicy.ts L29-40: allow-lossy.
    • The secret-carrying headers:
      • open-sse/utils/selfHop.ts L1-12
      • src/lib/guardrails/videoBridgeBrokerAuth.ts L7-19
      • open-sse/utils/requestLogger.ts L100 and L109-111: lease-owner is redacted.
    • Effort evidence:
      • open-sse/handlers/chatCore/attemptLogging.ts L513-516: providerRequest into pipelinePayloads. L575-582: requestBody from the client body.
      • src/lib/usage/callLogs.ts: L646-653 for the effort columns; L1047-1053 for the detail route.
      • src/lib/db/detailedLogs.ts L56-64: detailed logging.
      • open-sse/executors/codex.ts L1451-1465: the effort rewrite.

Evidence-class table

Claim Evidence class Command / result
The allowlist, the namespace rule, real-entry-point refusals and the tomllib refusal local integration test, fail-first The final tests on the pre-repair code (d19e3a69, the rebased 854803b0): FAILED (failures=6, errors=1). After the repair: tests.test_landscape_sweep_harness Ran 122 tests ... OK (skipped=3), again on the rebased head
Each new guard is needed local mutation check 10 named mutations, each targeted test run, file restored and sha256-checked: 10 caught. They cover each allowlist, each namespace check, MODEL_NAME, the value and type checks, the tomllib refusal, a job directory created before settings, and an echoed refused name
The one-slash model renders the same prompt as the control native, no model call (this round) The README's parity block as written, on a lane home staged by this branch for 20129. Both prompt-input commands exit 0 with 4 items each (3 developer, 1 user). cmp of the normalized files exits 0; neither output contains the model slug. Earlier coordinator run (installed profile, prompt argument): 5 against 5 items, identical after the same normalization, and the negative control sharedgw/cx/gpt-6-astra differed (3 items)
20129 serves the one-slash model at max historical live host observation (2026-09-27, this host) A live call returned 200 with reasoning. A probe without -m sent the profile's gpt-6-astra and got 401 six times, which is why -m is required
Registry tests, validation and the pre-push gate local and native git push 57 registry tests OK; scripts/validate.py passed after re-registration per docs/lanes.md. This push went through #438's gate: pre-push: running the registry tests on 74c233d7…, then OK

Review round (review-439; one review, one repair)

# Finding Disposition
1 Credential-carrying OmniRoute headers passed the word filter Fixed: allowlist of four switches in both entry points. Negative tests for self-hop, video-bridge-broker, lease-owner and session-id, through the builder, settings and codex_call.sh start
2 requestBody is inbound, not outbound Fixed after checking the cited lines: README and row 88 name pipelinePayloads.providerRequest and the detailed-logging condition. The column rule is narrowed
3 A . passed in the provider segment Fixed: the namespace rule and diagnostic in both entry points; my.gw/gpt-6-astra-max is tested
4 The no-partial-state test called the pure settings() Fixed: it runs codex_call.sh start (mutation M8 is caught)
5 "Failed first" was not reproducible from git Fixed: row 87 says main already refused a second slash, and the test failed first on the diagnostic and the dotted namespace
6 The README stated tomllib checks as unconditional Fixed: the runner fails closed with staged headers, and the README states the 3.9 behaviour
7 The parity check compared counts; its commands had no recorded run Fixed: content comparison, run as written (above)
8 The alias mechanism was unsourced Fixed: the mechanism is removed and the observation kept
9 Body errors (NAME:VALUE, base, correlation-id citation, paths) Fixed in this body

Residual risks

  • Without tomllib, a header-less lane still skips the comparison, so a hand-edited lane home that carries headers goes unchecked on Python 3.9.
  • The builder's read-back refusal (a table that does not parse back) has no test.
  • Headers stop at 20129. The effect of the lost Codex identity headers on 20128's cache hits is not measured.
  • A lane home under the temporary directory makes Codex 0.157.1 skip its apply_patch and sandbox PATH aliases (codex-rs/arg0/src/lib.rs L344-357). The parity run printed that warning. On Linux the sandbox falls back to the codex executable (L281-291), and the effect on shell-invoked apply_patch is not measured. This is a follow-up; it is out of scope for this round.
  • CI: validate-macos failed once on an unrelated codex_lane interruption-timing test, and the rerun passed. It is recorded as a flake to watch.
  • Harness acceptance on 20129 is still open (Status below).

Status

  • Built by: GPT-6 (gpt-6-astra, effort max), in two rounds, plus this coordinator repair round after review-439. The coordinator verified and committed.
  • Next, after merge:
    1. Restage the gateway lane home, which still holds the old two-slash model.
    2. Run a canary builder job on 20129 against a 20128 control, plus a multi-turn cache check. Record outbound effort from 20128's pipelinePayloads.providerRequest.
    3. Send the switch-on UTC time to peers for the evidence window.

🤖 Generated with Claude Code

@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Sep 27, 2026
Scout and others added 2 commits September 27, 2026 19:15
…ers and a one-slash model

The landscape-sweep GPT-6 lane can target the framework OmniRoute instance on
127.0.0.1:20129: all engines and output styles on, chained through node
sharedgw to the shared gateway on 20128. The shared 20128 stays unchanged
for the memory services and the other sessions.

- build_args.py --omniroute-header stages static provider headers
  (model_providers.<id>.http_headers). An example is
  x-omniroute-compression: allow-lossy. codex_job.py checks each job's headers
  against the lane home and records them per job.
- The lane's model has at most one provider segment: sharedgw/gpt-6-astra-max.
  Codex rust-v0.157.1 strips exactly one namespace segment when it looks up
  model metadata (models-manager/src/manager.rs L763-780). A two-slash slug
  such as sharedgw/cx/gpt-6-astra gets generic fallback metadata (model_info.rs
  L99-150), so both build_args and codex_job refuse it with a UsageError.
- README: the framework lane passes -m sharedgw/gpt-6-astra-max. Manual
  commands must pass -m, because under -p stack-worker the profile's model
  otherwise wins. The slashless alias is removed: the built-in codex provider
  claims gpt-6-astra* before stored aliases and answers 401. The
  prompt-input parity check is documented, and effort evidence is read from
  requestBody.reasoning.
- Two anti-pattern rows: the two-slash slug on a chained gateway, and a null
  call_logs effort column read as "effort not sent". No repository check
  enforces the second; the row says so.

Built by GPT-6 (gpt-6-astra, effort max) in two rounds: the header build, then
this amendment after verification found the two-slash fallback. The
coordinator verified and committed.

Checks, with TMPDIR outside every repository:
- The changed tests before the amendment: FAILED (failures=3).
- tests.test_landscape_sweep_harness and tests.test_codex_worker_lane:
  179 tests OK (skipped=11).
- The registry tests: 57 tests OK.
- codex debug prompt-input for sharedgw/gpt-6-astra-max and
  cx/gpt-6-astra-max: identical item counts. Base instructions are identical
  by the bundled-catalog lookup rule.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…utbound effort evidence)

One repair round for review-439's two major and seven minor findings.

- Static provider headers: an allowlist of OmniRoute 3.8.51's four per-request
  switches replaces the credential-word denylist. The denylist admitted
  x-omniroute-self-hop, -video-bridge-broker and -lease-owner, which carry
  secrets. The builder and runner never echo a refused name or value.
- Model names: the provider segment follows Codex's namespace rule, letters,
  digits, '_' and '-' only (manager.rs L763-780), in both entry points, with
  the fallback-metadata diagnostic. The no-partial-state test now runs the real
  `codex_call.sh start`.
- The runner refuses staged headers without tomllib; header-less lanes keep
  Python 3.9. The builder's read-back is unconditional (staging already needs
  3.11).
- Effort evidence: the outbound request is pipelinePayloads.providerRequest in
  the detailed call log. requestBody is the body the gateway received. A
  non-null reasoning_effort_upstream is positive evidence.
- README: a content-based parity check (without id and create_time, with the
  model masked), run as written; the alias mechanism is dropped and the 401
  observation kept.

The tests failed first (5 failures, 1 error) and now pass: 122 tests OK. A
mutation pass caught 10 of 10 guard mutations.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/gpt6-lane-fw-instance-20260927 branch from 854803b to 74c233d Compare September 27, 2026 23:16
@seathatflowsinourveins
seathatflowsinourveins merged commit b26add9 into main Sep 27, 2026
29 of 32 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the claude/gpt6-lane-fw-instance-20260927 branch September 27, 2026 23:47
seathatflowsinourveins added a commit that referenced this pull request Sep 28, 2026
…first recording; GPT-6 review agree) + dashboard gate row (#442)

* Capability-gate host receipts round 2 (at main eb5ca78), superseding the first recording, with GPT-6 review agree

The three gates re-ran live through host_receipts.py record --supersedes at the published
repair (#440). Each receipt quotes only its own verdict line, which carries the retained
results' sha256. The GPT-6 cross-family reviewer independently re-applied assertions.js to
the retained results (260/260, 10/10, 10/10 component verdicts; hashes 3/3; M13 sentinels
52/52) and recorded codex_lane agree on each receipt itself. The grand-dashboard gets the
codex-capability-gate row.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Register the round-2 capability-gate receipts and dashboard row (hot files)

docs/lanes.md hot-file protocol after merging origin/main at b26add9 (#439): main's
manifests/evidence.json with the three -2 receipts and observability/grand-dashboard/state.json
registered, and the component evidence matrix regenerated (flip-rule violations 0).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant