diff --git a/blueprints/gpt6-lane-compression-ab/PREREGISTRATION.md b/blueprints/gpt6-lane-compression-ab/PREREGISTRATION.md new file mode 100644 index 000000000..e2a04fb03 --- /dev/null +++ b/blueprints/gpt6-lane-compression-ab/PREREGISTRATION.md @@ -0,0 +1,474 @@ +# GPT-6 lane compression A/B preregistration + +**DRAFT — not sealed, not frozen, not run.** Repair round for draft PR #431 at +`f201e07b`, 2026-09-27, foundation lane. No execution or adoption authority. +The machine-readable contract is [preregistration.json](preregistration.json). +All model results, route qualification, pilot measurements, powered sample +sizes and confirmatory resource allocations remain `null`. + +This experiment compares all-attempt entry-gateway input plus output tokens, +subject to paired quality, exact-record and transport gates. Its ceiling is +**the six named task domains per role**. It cannot authorize moving a production +builder, reviewer or researcher role. Reviewer/researcher output-style conclusions +cover **JSON-only records**; prose verdicts, research narrative and evidence-writing +need representative frozen tasks in a new preregistration. + +Research used installed search-first inline, find-skills discovery, TDD and +verification-before-completion at the user-specified document/test seam. The +installed Codex reports 0.157.1 and its prompt-input help describes a renderer. +No model was called, including no prompt-input invocation in this repair. +Seventeen installed OmniRoute source hashes were recomputed: all 17 match the +retained pinned upstream bindings. This is source identity, not live acceptance. +Direct shell gh failed to connect; read-only gh API through Context Mode supplied +release notes and pinned source. No installed skills/tools or host configuration +were changed. Sources and the exact review dispositions appear below and in JSON. + +The historical original authoring evidence is retained in JSON. This round uses +the dated checks/findings/remaining-gates structure from +[PR #416 build evidence](https://github.com/seathatflowsinourveins/native-agent-stack/blob/b6f36d8cac660b47cafc827ee4595cf0bc74459a/blueprints/compaction-window-ab/build-evidence.json). +That file was absent in this worktree and read at the cited upstream revision. + +## Runner and native configuration + +**Choose Harbor v0.23.0 conditionally, under option (b).** Harbor at `1e5c5c6d` +passes only the final slash-separated segment to `--model`. Main `3c823808` does +the same; no release after v0.23.0 was found in the releases API checked on +2026-09-27. The canonical models below are therefore **not what Harbor currently +emits**. The emitted argument for every cell is `gpt-6-astra-max`. A working +20129-side route for that argument has not been identified or qualified. +The established built-in `gpt-6-astra*` precedence rules out assuming an ordinary +stored alias fixes it. This is an open sealing gate, not a fabricated mapping. +[Harbor command](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1339-L1449), [inspected main](https://github.com/harbor-framework/harbor/blob/3c82380859d187957cfd5cd64802b076d9779550/src/harbor/agents/installed/codex.py#L1502-L1605). + +| Alternative | Source-backed comparison | +| --- | --- | +| Later Harbor release/main | Release query returned v0.23.0 as latest; inspected main still strips prefixes. An upgrade alone is not a demonstrated repair. | +| Inspect inspect_swe 0.2.71 | Resolves a catalog slug into `--model` and routes through its own OpenAI bridge. Profile, wire and correlation equivalence require qualification. [Command and bridge](https://github.com/meridianlabs-ai/inspect_swe/blob/7eb8dd64309db4cd0f6bdf1d0ffd9786a74a4088/src/inspect_swe/_codex_cli/codex_cli.py#L483-L637). | +| promptfoo Codex SDK 0.123.1 | `config.model` passes intact into thread options; Codex SDK passes it intact to `--model`. Explicit thread resume exists, while deep tracing forces fresh threads. This is the strongest alternate for prefix preservation; native task/verifier/container integration is unqualified. [Provider](https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/providers/openai/codex-sdk.ts#L1034-L1134), [SDK](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/sdk/typescript/src/exec.ts#L91-L178). | + +Harbor retains the existing native task/verifier/multi-step route without a new +runner implementation. **Overturn this choice** before confirmation if promptfoo's +Codex SDK passes the same merged-config, prompt-input, three-turn tool/verifier, +correlation and both-hop effort checks while Harbor has no supported route or +has unequal native semantics. Amend once before data; never pool runners. + +Route qualification belongs to the coordinator. Compare +`codex debug prompt-input` for the slashless candidate against the canonical +one-slash slug with identical semantic settings. Require the expected five items +and equivalent template, tools, Responses Lite and multi-agent metadata after +declared identity/path normalization; item count alone is insufficient. Then +require one attributable live 200 at max through 20129 and prove C's emitted +bare route equivalent to `cx/gpt-6-astra-max`. Reuse the user's established +one-namespace metadata fact; do not re-probe it during repair. + +Adopt the **native semantic deep merge** of frozen `config.toml` followed by +`stack-worker.config.toml`, with no `profile` or `profiles` keys. Hash both inputs, +merged TOML, uploaded config and resolved effective settings. Codex's merge has +normalization and replacement rules beyond a shallow dictionary update. Harbor +uploads this file via `config`, then applies its runtime/MCP overrides. +[Profile loader](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/config/src/loader/mod.rs#L286-L340), [merge implementation](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/config/src/merge.rs#L56-L185), +[Harbor upload](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1207-L1296). + +Equivalence gate: same binary/model/cwd/project/system layers and effective +forced flags, compare **resolved configuration and prompt-input** from the +`-p stack-worker` base-plus-profile reference to the merged-file run without `-p`. +Repeat after Harbor upload. Configuration provenance paths may differ; semantic +settings, prompt content and tool definitions must match. Freeze the native +resolved-config inspection command before qualification; it was not run here. + +Harbor's actual command contract is: + +```text +codex exec [resume --last on turns 2/3] + --dangerously-bypass-approvals-and-sandbox --skip-git-repo-check + --model gpt-6-astra-max --json --enable unified_exec + -c model_reasoning_effort=max -- +``` + +The upstream wrapper optionally sources NVM, redirects `2>&1 ` at each joined hop and inspect the +stored client/provider request body's **`reasoning.effort`**, printing only that +field. Freeze actual detail selectors during qualification. Null effort columns +are expected without encrypted reasoning and **never prove drift**; they only +corroborate. Missing body evidence stays unknown, observed non-max is drift. +[Conditional columns](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L628-L653), [detail API](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/app/api/usage/call-logs/[id]/route.ts#L1-L22). + +| Quantity | Boundary | +| --- | --- | +| Uncached input | `sum(tokens_in - tokens_cache_read)` | +| Total positions | `sum(tokens_in + tokens_out)` at the entry only | +| Weighted input | uncached input + cache ratio × cached input | +| Weighted total | weighted input + output ratio × total output | +| Reasoning | Output subset, reported separately, never added again | +| Compression diagnostic | Joined 20129 `tokens_compressed`; includes reactive compaction, not billed savings | + +Exclusive-window analytics can corroborate compression diagnostics but are +optional and never additive with per-call savings. Retain failures, retries and +cancellations. Null usage remains null, with a lower/upper interval only if +supported by stored nonsecret request size **and** provider serialization/context +and output caps. Request bytes alone cannot bound total tokens; pilot maxima +are not hard caps. Without a defensible upper bound, cost is inconclusive. +Require the winner to remain cheaper under candidate upper/control lower totals +and all sensitivity weights. This replaces automatic rejection of every null +499 row while preserving uncertainty. [Nullable counters](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107). + +Actual subscription prices remain unknown. Preregister sensitivity assumptions +**cache ratio [0,1], output ratio [1,10]**, with uncached input as unit weight. +Require the economic guard at all four corners for whole-session and turn 1/2/3 +strata; test weighted candidate minus 1.01×C using simultaneous paired bounds. +The weighted inequality is linear in weights, so the corner checks cover the +rectangle. Report crossover weights. This is robustness to stated assumptions, +not GPT-6 dollar pricing or a measured subscription-limit meter. + +`x-omniroute-connection` is a **connection pin**, not session affinity. Affinity +uses `x-codex-session-id`, `x-session-id`, `x-omniroute-session`, then body session +identifiers, `prompt_cache_key`, then first-input hash. Live zone instead uses +`x-omniroute-session-id` or a generated body/provider/connection fingerprint. +Trace cache keys at both hops, salted accounts and per-arm account spread/cache +read rates. The review's account-collapse risk remains unobserved; verify it in +qualification. [Affinity](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/services/sessionAffinityPin.ts#L197-L221), [live-zone fallback](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/sessionManager.ts#L103-L145). + +## Pilot, statistics and resources + +Run no confirmation until an **excluded pre-confirmatory qualification pilot** +has measured each cell. Proposed pilot: 12 paired draws per role, two per task, +all five cells, 180 native three-turn sessions, zero primes. Its provisional +24M-token/seven-day/30-minute-session ceilings are monitored planning limits, +not evidence it fits or execution authority. Preflight native request/output +bounds and reserve before starting. An incomplete pilot cannot size confirmation. +Keep all its usage, wall time, failures and refusals. + +Retain joint five-arm final success/unrecovered-failure patterns and per-cell +tokens, wall time, cache by turn, account spread and null-usage bounds. Then use +the pinned NumPy/SciPy algorithms in JSON to simulate power and margin-null +calibration. Candidate n grid: 120, 240, 480, 960, 1920, 3840; 10,000 simulations +per law, seed 20260929. The random primitives are pinned to +[NumPy 2.4.0](https://github.com/numpy/numpy/blob/c5ab79c14c98bfda1e60770ffa23a6130f8267b7/numpy/random/_generator.pyx#L299-L310). Preserve empirical joint dependence and a declared +sensitivity grid. For harmless-arm power, permute all five arm labels within +a resampled pilot draw, preserving joint outcomes with equal marginal rates. +Sensitivity laws use shared-or-independent uniform draws with fixed success and +unrecovered-failure thresholds; exact laws and feasibility rules are in JSON. Choose the smallest n whose lower 95% exact power bound is at least .80 +for jointly qualifying four harmless candidates per role, and whose upper +family false-promotion bound at margin-null is at most .055 (Monte Carlo +tolerance .005 around .05). If none qualifies, keep the gate open. No statistical +simulation or pilot result is claimed by this repair. + +Primary success harm is `mean(C_success - candidate_success)`; failure harm is +`mean(candidate_unrecovered_failure - C_unrecovered_failure)`. Both margins are +five percentage points. Use `scipy.stats.bootstrap(paired=True, method="percentile")`, +99,999 whole-draw resamples, `alternative="less"`, fixed seed 20260927. Invert the one-sided upper +percentile interval against +.05. This is approximate and must pass the pre-data +calibration; a nonsignificant equality test is not non-inferiority. +[SciPy bootstrap](https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_resampling.py#L300-L394). Tango's paired-proportions score interval is a +source-backed alternative, recorded but not implemented with a new solver. +[Tango 1998](https://doi.org/10.1002/(SICI)1097-0258(19980430)17:8%3C891::AID-SIM780%3E3.0.CO;2-B). + +Exact fallback when fewer than 20 discordant draws, degenerate or nonfinite +bootstrap: h counts harmful discordance and b counts beneficial discordance. +At test level a use one-sided Clopper-Pearson `U(h,n,1-a/2) - L(b,n,1-a/2)`. +This union-bound upper limit covers **net harm**, without independence between +h and b. Require it below .05; invert monotonically for p. All-concordant data +have a nonzero exact uncertainty bound. Unsupported results get p=1. +[SciPy exact interval](https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_binomtest.py#L52-L115). + +Combine the two required endpoints with candidate `p=max(p_success,p_failure)` +(intersection-union), then Holm across **12 role-candidate hypotheses** at .05; +missing/unrun hypotheses get p=1. This avoids treating each mandatory component +as a separate promotion claim. [Intersection-union reference](https://doi.org/10.1214/ss/1032280304), +[Holm implementation](https://github.com/statsmodels/statsmodels/blob/278ff9950636cdd4939b4055e339a8e681d79cab/statsmodels/stats/multitest.py#L99-L149). Control and candidate still need observed +combined success ≥.90, at least one success for each task, zero exact-record +corruption and valid native controls. + +Confirmation uses independent uniform draws over each role's six tasks, paired +across all five cells. Retain the Williams orders and their reversals recorded +in JSON. Freeze realized schedule after the pilot and before confirmatory data; +no outcome-driven balancing, replacement or extra sampling. There are **zero +priming sessions**, no cold/warm labels and no cache flush. Analyze measured +cache state by user turn 1, 2 and 3, keeping all requests/retries attached to the +session. Fresh per-arm/draw identities do not prove a cold cache. + +Sample size, confirmation token cap and wall allowance remain null until sizing. +Use pilot per-cell one-sided upper mean cost/wall estimates ×1.25 and the larger +99% resampled aggregate schedule estimate. Add pilot/control/exploration costs, +setup/teardown and a separately qualified inflight/cancellation reserve. With +one session at a time, sum all five cells' wall time. A statistical forecast +does not create a hard provider charge cap. Never reduce powered n to fit an +insufficient resource envelope. The old 5,400-session plan allowed only 3,703.70 +tokens and 112 seconds per session; that arithmetic is verified, but the cited +13,806-token request came from a different invocation and was not a measured +minimum for this cohort. + +Pilot-calibrate routine error/cancellation thresholds over rolling 20-session +blocks using conservative exact rate bounds and a .01 whole-run false-stop +allocation. Freeze finalization/reconciliation delays from pilot timing. Immediate +stops remain corruption, authentication/quota refusal, observed model/effort/config +drift, failed controls, ambiguous joins and insufficient resource reserve. +Null effort columns and recovered transient failures are not drift. + +Use native Harbor cancellation, then only the owned recorded process group with +bounded interrupt/TERM/KILL escalation; verify local termination and reconcile +entry-gateway rows. Never stop gateways or unrelated sessions. Restore only the +owned 20129 settings snapshot. Unknown remote cancellation/trailing usage stays +unknown. None of this lifecycle work executes while drafting. + +After quality gates, rank complete all-attempt total tokens, with the preregistered +1% tie preference C, D0, D1, A0, A1 and evidence of positive paired saving. Economic +guards use simultaneous bounds across candidates, rectangle corners and turn/ +session strata. Incomplete, underpowered or weight-sensitive results retain C. + +## Review repair dispositions + +“Fixed by gating” means the protocol defect is repaired but an explicitly open +qualification result prevents sealing. It is not passed host acceptance. +PLAUSIBLE findings remain marked as unobserved where source inspection cannot +establish their occurrence. No whole finding was declined; the unsupported +request-size-only total usage bound in M9 was specifically declined. + +| Finding | Disposition | Repair, verification and source | +| --- | --- | --- | +| B1 | fixed by gating | Canonical hop names fixed; retain Harbor option (b) with unresolved supported route, full prompt equivalence and live-200-max gates; compare three alternatives. Verified stripping in v0.23.0/main. Established namespace/route facts reused. PLAUSIBLE bare-name 401 not re-probed; no working route asserted. [Source 1](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1339-L1449) [Source 2](https://github.com/harbor-framework/harbor/blob/3c82380859d187957cfd5cd64802b076d9779550/src/harbor/agents/installed/codex.py#L1502-L1605) [Source 3](https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/providers/openai/codex-sdk.ts#L1034-L1134) [Source 4](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/sdk/typescript/src/exec.ts#L91-L178) [Source 5](https://github.com/meridianlabs-ai/inspect_swe/blob/7eb8dd64309db4cd0f6bdf1d0ffd9786a74a4088/src/inspect_swe/_codex_cli/codex_cli.py#L483-L637) | +| B2 | fixed by gating | Merged semantic config, actual forced command, equal-arm deviations, host-network overlay/probe, admin containment gate and native [[steps]] mapping. Read loader/merge, Harbor upload/command/compose/multi-step code. WSL2 reachability and config equivalence unobserved, gated. [Source 1](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/config/src/loader/mod.rs#L286-L340) [Source 2](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/config/src/merge.rs#L56-L185) [Source 3](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1339-L1449) [Source 4](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/environments/docker/docker.py#L350-L420) [Source 5](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/trial/multi_step.py#L25-L110) | +| B3 | fixed by gating | Pre-confirmatory pilot measures all cell costs/wall times; powered n, confirmation budget/reserve/wall fields now null pending sizing. Zero primes. Old arithmetic verified: 5,400 sessions, 3,703.70 tokens/session, 112 s/session. Historical 13,806-token row is a different invocation, not a measured lower bound for this cohort; no new host receipt read. [Source 1](https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_resampling.py#L300-L394) [Source 2](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/trial/multi_step.py#L25-L110) | +| B4 | fixed by gating | Usage at entry only joined on X-Correlation-Id; upstream-hook observer specified; no hop summation. Request-body effort at both hops; nullable columns corroborate only. Verified call_logs schema/conditional effort and header emission; native Codex header export not found in inspected SSE path. Observer and second-hop exact join remain unqualified. [Source 1](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107) [Source 2](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/handlers/chatHelpers.ts#L1172-L1184) [Source 3](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L628-L653) [Source 4](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/app/api/usage/call-logs/[id]/route.ts#L1-L22) [Source 5](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/codex-api/src/sse/responses.rs#L30-L101) [Source 6](https://github.com/mitmproxy/mitmproxy/blob/6c09d56e4c29a92f5ad01b03199977584b8ea14f/mitmproxy/proxy/layers/http/_hooks.py#L7-L39) | +| B5 | fixed by gating | Paired net-harm bootstrap with explicit conservative exact fallback; 12 candidate intersection-union Holm tests, unrecovered failures, pilot power and null calibration. Recomputed reviewer's old-gate probabilities .003102/.306404/.784372; cited pinned bootstrap/exact/Holm code. No power result claimed without pilot. [Source 1](https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_resampling.py#L300-L394) [Source 2](https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_binomtest.py#L52-L115) [Source 3](https://github.com/statsmodels/statsmodels/blob/278ff9950636cdd4939b4055e339a8e681d79cab/statsmodels/stats/multitest.py#L99-L149) [Source 4](https://doi.org/10.1002/(SICI)1097-0258(19980430)17:8%3C891::AID-SIM780%3E3.0.CO;2-B) [Source 5](https://doi.org/10.1214/ss/1032280304) [Source 6](https://github.com/numpy/numpy/blob/c5ab79c14c98bfda1e60770ffa23a6130f8267b7/numpy/random/_generator.pyx#L299-L310). | +| M1 | fixed by gating | Registered-key D/A cells and TTL 60 are confirmatory; key name through env_key/extra_env templates only; reuse acceptance gated. Established host configuration reused; inspected principal/session and environment-template source, no key file access. [Source 1](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/liveZone.ts#L127-L132) [Source 2](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449) [Source 3](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/base.py#L560-L648) [Source 4](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/utils/env.py#L4-L65) | +| M2 | fixed | Sensitivity cache ratio [0,1], output ratio [1,10]; robust all-corner/turn guard without claiming actual prices or subscription meter. These ranges are explicit protocol assumptions over source-defined token components; actual price fields remain null. [Source 1](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107) [Source 2](https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_resampling.py#L300-L394) | +| M3 | fixed by gating | Four native output-shape canaries with whole-string eligibility and native-reachability results; shell minifier nonreachability cannot count as engine acceptance. Verified shell/unified_exec headers and whole-string JSON.parse. MCP/custom shapes and actual engine application remain qualification results. [Source 1](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/codexResponses/index.ts#L137-L145) [Source 2](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/tools/context.rs#L524-L601) [Source 3](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/tools/mod.rs#L97-L124) [Source 4](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/bodyAdapter.ts#L116-L145) | +| M4 | fixed by gating | Direct Chat key-2 collision stimulus rebuilt; assert multipart pin part unchanged; native Responses reachability and wire format independently gated. Source key construction/owner ordering/replacement verified. input_text bypass is source-supported concern; no native collision observed. [Source 1](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/session-dedup/index.ts#L291-L345) [Source 2](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/bodyAdapter.ts#L116-L145) [Source 3](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/strategySelector.ts#L367-L386) | +| M5 | fixed | Reviewer/researcher style conclusions restricted to JSON-only records; no prose-verdict/evidence harmlessness claim. All 12 task schemas inspected; style catalog targets prose/code. PLAUSIBLE effect magnitude remains unmeasured; narrowed scope uses reviewer's alternative fix. [Source 1](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/outputStyles/catalog.ts#L32-L194) [Source 2](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449) | +| M6 | fixed by gating | Connection pin separated from session affinity; prompt_cache_key at both hops, account spread/cache reads by arm, prior false correction repaired. Header/body/fallback precedence verified. PLAUSIBLE account collapse not observed and not claimed; qualification required. [Source 1](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/services/sessionAffinityPin.ts#L197-L221) [Source 2](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/executors/codex.ts#L1498-L1573) | +| M7 | fixed | Every session including C scores exact values; application/reachability is separate engine coverage, never exclusion or vacuous pass. Source allows no-op/ineligible engine paths; public document contract now explicitly defines both outcomes. [Source 1](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/codexResponses/index.ts#L137-L145) [Source 2](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/strategySelector.ts#L367-L386) | +| M8 | fixed | Drop all separate priming and cold/warm labels; stratify measured cache analysis by turn. PLAUSIBLE lack of warming not established experimentally; native session/affinity source supports avoiding an unverified priming benefit. No cache-effect claim. [Source 1](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/services/sessionAffinityPin.ts#L197-L221) [Source 2](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/executors/codex.ts#L1498-L1573) | +| M9 | fixed by gating | Pilot-calibrated operational stops; null usage intervals and worst-case ranking. Request size alone is declined as a bound on total input+output usage. Nullable input/output/reasoning schema verified. Routine error prevalence unmeasured. Require supported output/serialization caps as well as stored request size; without them cost gate stays open. [Source 1](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107) [Source 2](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L628-L653) [Source 3](https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_binomtest.py#L52-L115) | +| M10 | fixed | Explicit ceiling: named synthetic task domains only, no production role migration or prose-record adoption. Six-task packets and four builder smoke cases inspected; narrower scope states the design's limitation without adding representative tasks. [Source 1](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/verifier/verifier.py#L165-L250) [Source 2](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/outputStyles/catalog.ts#L32-L194) | +| m1 | fixed | Replace Nine with Seventeen installed source files in correction log. Recomputed all 17 retained installed SHA256 bindings: 17 matched. No new live acceptance implied. [Source 1](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107) [Source 2](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/codexResponses/index.ts#L137-L145) | +| m2 | fixed | Per-call tokens_compressed primary diagnostic; includes reactive compaction, not billed savings. Analytics optional, never additive. Inspected callLogs L107 and chatCore L1916-1919/L2164; source-backed attribution improvement. [Source 1](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107) [Source 2](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449) | +| m3 | fixed | promptfoo gateway example uses sharedgw/gpt-6-astra-max; Codex SDK alternate preserves the same one-slash slug. Verified provider model propagation and SDK --model code path; no promptfoo execution. [Source 1](https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/providers/openai/responses.ts#L1195-L1210) [Source 2](https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/providers/openai/codex-sdk.ts#L1034-L1134) [Source 3](https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/sdk/typescript/src/exec.ts#L91-L178) | +| m4 | fixed by gating | Require ambient OPENAI_API_KEY/auth-copy switches unset and use OMNIROUTE_FW_API_KEY template; verify no literal/partial secret logs. Harbor can inject OpenAI credentials and KEY-sensitive env handling verified. Presence checks and redaction runtime acceptance left to coordinator; no values opened. [Source 1](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1339-L1449) [Source 2](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/base.py#L560-L648) [Source 3](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/utils/env.py#L4-L65) | +| m5 | fixed by gating | Responses->Responses sharedgw format specified; no Chat fallback; real encrypted/custom-item preservation gated. Adapter and pre/post-translation stage source inspected; configured/actual node format still unobserved. [Source 1](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449) [Source 2](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/bodyAdapter.ts#L116-L145) [Source 3](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/strategySelector.ts#L367-L386) | +| m6 | fixed | Live-zone session fallback derived from request body/provider/connection is explicit and distinct from affinity extraction. Inspected chatCore L1868-1881 and sessionManager L103-145. [Source 1](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/sessionManager.ts#L103-L145) [Source 2](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449) | +| m7 | fixed | Tests now enforce canonical model hops and entry gateway ledger, plus all structurally checkable blockers. Before document repair: requested unittest command exit 1, 10 tests, 12 failures, including old L61/L122 expectations. [Source 1](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1339-L1449) [Source 2](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/handlers/chatHelpers.ts#L1172-L1184) [Source 3](https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107) | +| m8 | fixed | Exploratory repetition unit explicit; keyed exploratory cell removed; original two builder steps are turns 1/2 and record step is turn 3. Pinned task.toml contains exactly create-file and append-content; native resume implementation inspected. Wrapper remains local integration, not unchanged upstream execution. [Source 1](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/examples/tasks/hello-multi-step-simple/task.toml#L15-L31) [Source 2](https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/trial/multi_step.py#L25-L110) | + +## Open sealing gates + +- **runner-route**: Supported 20129 slashless route preserving canonical native metadata, five-item/full prompt equivalence and coordinator live 200 at max; C route equivalence too. No mapping has been found/qualified. +- **merged-profile**: Materialize/hash native semantic base+stack-worker merge; no profile/profiles keys; same resolved config and prompt-input versus -p reference before/after Harbor forced flags. +- **network-containment**: Hash extra_docker_compose overlay; prove actual container reachability and separately qualify upstream-supported containment of passwordless admin APIs; trusted-task constraints alone are not a security boundary. +- **header-capture**: Materialize/hash upstream-hook response-header observer, qualify privacy/SSE/cache semantics and X-Correlation-Id entry joins including failures; demonstrate a separate exact second-hop effort join. +- **effort-detail**: Qualify detail-API selectors and inspect only reasoning.effort at joined hops; null corroboration columns are allowed, missing body stays unknown. +- **two-hop**: Freeze Responses wire format and prove three native steps, genuine call IDs/encrypted reasoning replay, prompt_cache_key, session/account affinity and arm account spread/cache rate. +- **effective-config**: Read back C off/exclusions; keyed D/A engine/style plans, 60-minute live zone, optional dependencies and safety compaction. Existing key/TTL are facts, reuse is not accepted. +- **native-harnesses**: Harbor 0.23.0, Codex 0.157.1, chosen observer, statistics, actual task dependencies/images and exact build/config hashes through supported upstream commands. +- **native-controls**: Materialize/hash three-turn wrappers and graders; native known-pass/fail/malformed controls, four numeric output shapes, rebuilt Chat collision plus separately observed native reachability. +- **pilot-power-resources**: Complete excluded pre-confirmatory pilot; calibrate NI at margin, simulate power and operational stops; freeze powered n, complete token budget, reserve, wall time and schedule. +- **usage-sensitivity**: Confirm usage fields and finalization, missing-row input/output bounds, all-corner sensitivity and simultaneous cost bounds; actual dollar prices remain unknown. + +Qualification inputs/ceilings must be frozen by the coordinator before live +qualification; confirmation inputs/schedule/resources only after excluded pilot +results. This document freezes neither. All open gate results are null, and +material changes after confirmatory data require a new cohort. Wider role/prose +adoption is outside this experiment's ceiling, not a gate this cohort can close. + +## Local structural verification + +The new tests ran before changing either protocol file. The requested command +with `PYTHONDONTWRITEBYTECODE=1`, `TMPDIR=/var/tmp/claude-431` and +`GIT_OPTIONAL_LOCKS=0` returned **exit 1**: `Ran 10 tests in 0.009s`, +`FAILED (failures=12)`. It rejected the old model/usage assertions and all missing +blocker contracts. Full sanitized returned output and subsequent acceptance +commands are retained in `repair_round_20260927.checks` in JSON. + +After repair, the same contract command returned **exit 0**: +`Ran 10 tests in 0.011s`, `OK`. + +The required broader command +`python3 -m unittest tests.test_osv_lockfile_coverage tests.test_blind_checkout tests.test_workflow_security_coverage` +returned **exit 1**: `Ran 57 tests in 24.869s`, `FAILED (errors=11)`. +All eleven errors are blind-export ancestor refusals: the requested +`/var/tmp/claude-431` also has an existing read-only `.git` directory. The guard +at `tools/sota-convergence/blind_checkout.py:970-973` checks its presence even +when empty. No metadata was removed and no test was weakened. The separate +`tests.test_blind_checkout.RepositoryClassificationTests -v` command returned +**exit 0**, `Ran 2 tests in 0.151s`, `OK`; there are no new unclassified strings. + +`python3 scripts/validate.py` returned **exit 0**: +`{"components": 69, "hashed_files": 7348, "profiles": 4, "receipts": 159, "status": "passed"}` +and `Integrity and scope checks only; no live provider or GPU execution.` +`git diff --check` returned **exit 0** with no output. Every Python command used +the requested TMPDIR and bytecode suppression. Final post-record checks are +appended in the JSON round record; the failed broader run remains retained. + +These are local structural/integrity checks, not Harbor upstream acceptance, +gateway probes, model trials or statistical power evidence. The coordinator +owns any hash re-registration and commit; no git metadata or evidence manifest +is written by this repair. Corrections, including the earlier false affinity +correction and the seventeen-file count, remain in the draft's anti-pattern log. diff --git a/blueprints/gpt6-lane-compression-ab/preregistration.json b/blueprints/gpt6-lane-compression-ab/preregistration.json new file mode 100644 index 000000000..ca6020bcb --- /dev/null +++ b/blueprints/gpt6-lane-compression-ab/preregistration.json @@ -0,0 +1,3936 @@ +{ + "schema_version": 1, + "status": "DRAFT", + "frozen": false, + "run_started": false, + "execution_authorized": false, + "created": "2026-09-27", + "lane": "foundation", + "base_revision": "0f76651d", + "objective": "Choose GPT-6 compression configurations per frozen role domain using upstream grading, exact-value preservation and complete two-hop accounting.", + "results": null, + "research": { + "method": "Installed sources/help where observational, pinned changelogs and sources via gh api, then source-based design. search-first researcher and installed find-skills/tdd/verification skills; no installs.", + "structural_model": "#416 blueprints/compaction-window-ab, read-only sibling draft, structure only; no outcomes or thresholds copied as evidence.", + "local_test_reference": "tests/test_token_e2e_preregistration.py at 0f76651d", + "mapping": [ + "catalogs/convergence-practice/source-review.json", + "catalogs/convergence-practice/architecture-wave/harbor-framework__harbor.json", + "catalogs/sota-convergence/manifest-20260926.json", + "manifests/stack.json" + ], + "discovery_limit": "Skills discovery used installed skill sources. Web open was unsupported and a search returned no usable output; no catalog-wide absence claim. Direct shell gh could not connect; network-capable research tool gh api succeeded.", + "source_date": "2026-09-27", + "repair_method": "2026-09-27: installed Codex version/help, installed skill discovery and OmniRoute source; gh api release/tag/source reads through Context Mode after direct shell network failure; inline only, no agents or installations. No Harbor release after v0.23.0 found in releases API; inspected main 3c823808 still strips prefixes. Promptfoo/inspect_swe alternatives checked at their latest returned releases.", + "repair_skills": [ + "search-first (inline)", + "find-skills (installed discovery only)", + "tdd (user-authorized document seam)", + "verification-before-completion", + "openai-docs (user-required installed/tag source order)" + ], + "pattern_source": "https://github.com/seathatflowsinourveins/native-agent-stack/blob/b6f36d8cac660b47cafc827ee4595cf0bc74459a/blueprints/compaction-window-ab/build-evidence.json", + "scope": "Repair of draft f201e07b only. Historical original authoring records remain historical. No runtime/pilot/route acceptance in this round.", + "policy_constants": "Sensitivity ranges, .80 power target, .055 calibration tolerance, n grid, 1.25 sizing factor and pilot ceiling are explicit proposed protocol assumptions. Upstream citations support statistical/runner primitives, not claims these constants were measured or optimized." + }, + "sources": [ + { + "id": "or-live-zone", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/liveZone.ts", + "lines": [ + 127, + 132 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/liveZone.ts#L127-L132", + "claim": "Principal and session required; absent context calls compress(body) at 376-395.", + "installed_build": "dd6e9607e", + "installed_file_sha256": "e959421584a9a2523a2615f4892dd37ce43e8add67fc3972c699f766371798da", + "gh_api_exit": 0, + "installed_equals_upstream": true, + "evidence_class": "read-only installed source and gh-api byte comparison" + }, + { + "id": "or-chat-core", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/handlers/chatCore.ts", + "lines": [ + 1425, + 1449 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449", + "claim": "Opt-out leaves reactive passes enabled. Styles 1689-1734; principal 1793; live zone 1866-1903; safety passes 2134-2160 and 2210-2231.", + "installed_build": "dd6e9607e", + "installed_file_sha256": "d244f5e9045e320161617d8b1488dfbe40f78dea74deb1aef73ae67bd7d534a9", + "gh_api_exit": 0, + "installed_equals_upstream": true, + "evidence_class": "read-only installed source and gh-api byte comparison" + }, + { + "id": "or-lite", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/lite.ts", + "lines": [ + 148, + 168 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/lite.ts#L148-L168", + "claim": "Default limit 2000 at line 24; truncates string tool content; default apply path 259-262.", + "installed_build": "dd6e9607e", + "installed_file_sha256": "255644fe84ae8e80e196c5bc25ce3dea1321c4ade274059a1c635721f722b29c", + "gh_api_exit": 0, + "installed_equals_upstream": true, + "evidence_class": "read-only installed source and gh-api byte comparison" + }, + { + "id": "or-adapter", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/bodyAdapter.ts", + "lines": [ + 116, + 145 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/bodyAdapter.ts#L116-L145", + "claim": "Responses function_call_output becomes role tool content; restore 205-221,315-365; eligibility metadata is separate.", + "installed_build": "dd6e9607e", + "installed_file_sha256": "b2a836ff12f870e5d464f7b228e7b7a18a0a43fb72c54ff16bb70d94d80cdc01", + "gh_api_exit": 0, + "installed_equals_upstream": true, + "evidence_class": "read-only installed source and gh-api byte comparison" + }, + { + "id": "or-strategy", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/strategySelector.ts", + "lines": [ + 367, + 386 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/strategySelector.ts#L367-L386", + "claim": "Lite and stacked paths use Responses adapter; lite does not consult codex-responses eligibility metadata.", + "installed_build": "dd6e9607e", + "installed_file_sha256": "76a671f9af9b2be846854e1254c8cd236b8d67449cb93927199c579ab85c117a", + "gh_api_exit": 0, + "installed_equals_upstream": true, + "evidence_class": "read-only installed source and gh-api byte comparison" + }, + { + "id": "or-minify", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/engines/codexResponses/index.ts", + "lines": [ + 137, + 145 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/codexResponses/index.ts#L137-L145", + "claim": "JSON.parse then JSON.stringify; candidate selection also has size thresholds at 216-226.", + "installed_build": "dd6e9607e", + "installed_file_sha256": "8b53cc348e643c7b96a89ef496cb2d700ddeb36ceca948c2833666f29daafc20", + "gh_api_exit": 0, + "installed_equals_upstream": true, + "evidence_class": "read-only installed source and gh-api byte comparison" + }, + { + "id": "or-dedup", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/engines/session-dedup/index.ts", + "lines": [ + 291, + 345 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/session-dedup/index.ts#L291-L345", + "claim": "String key i collides with multipart key i*100000+p+1, and replacement lookup uses both key spaces.", + "installed_build": "dd6e9607e", + "installed_file_sha256": "00e5e81520fb4189256924cbfc4bd0800134ca525d5d778155db0eb9e64be64b", + "gh_api_exit": 0, + "installed_equals_upstream": true, + "evidence_class": "read-only installed source and gh-api byte comparison" + }, + { + "id": "or-engine-labels", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/engineCatalog.ts", + "lines": [ + 31, + 108 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engineCatalog.ts#L31-L108", + "claim": "Four engines carry lossy:false labels; this is classification, not exact-value acceptance.", + "installed_build": "dd6e9607e", + "installed_file_sha256": "d0029c910f3587cc352347ae1d6461caef8f758b0210f6d5ba4d95cd162ef04b", + "gh_api_exit": 0, + "installed_equals_upstream": true, + "evidence_class": "read-only installed source and gh-api byte comparison" + }, + { + "id": "or-styles", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/outputStyles/backCompat.ts", + "lines": [ + 13, + 28 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/outputStyles/backCompat.ts#L13-L28", + "claim": "Explicit nonempty styles override legacy caveman output mode; empty styles fall back unless legacy disabled.", + "installed_build": "dd6e9607e", + "installed_file_sha256": "7a2d5d004e4b028b97f7b59278c1b49fe2a1848e20337d803e6d51452271accd", + "gh_api_exit": 0, + "installed_equals_upstream": true, + "evidence_class": "read-only installed source and gh-api byte comparison" + }, + { + "id": "or-lossy-policy", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/lossyRequestPolicy.ts", + "lines": [ + 29, + 67 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/lossyRequestPolicy.ts#L29-L67", + "claim": "allow-lossy keeps plan; headerless stacked plan retains isSafeDefault steps.", + "installed_build": "dd6e9607e", + "evidence_class": "read-only installed source and gh-api byte comparison", + "live_acceptance": null, + "installed_file_sha256": "8831e43712b6b2a994a418022e548f31b8eebc46b9f6fa8c3a32e60f2eac25c5", + "gh_api_exit": 0, + "installed_equals_upstream": true + }, + { + "id": "or-caveman", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/caveman.ts", + "lines": [ + 403, + 445 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/caveman.ts#L403-L445", + "claim": "Rule rewrites applied to message content.", + "installed_build": "dd6e9607e", + "evidence_class": "read-only installed source and gh-api byte comparison", + "live_acceptance": null, + "installed_file_sha256": "55bbdc6bebf75d958d8f25dff391e1cbdcffb1c9845a34446e714d465af992d0", + "gh_api_exit": 0, + "installed_equals_upstream": true + }, + { + "id": "or-relevance", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/engines/relevance/index.ts", + "lines": [ + 36, + 84 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/relevance/index.ts#L36-L84", + "claim": "Sentence selection retains a subset in original order.", + "installed_build": "dd6e9607e", + "evidence_class": "read-only installed source and gh-api byte comparison", + "live_acceptance": null, + "installed_file_sha256": "94d1e363b74564df40b719aeeb89af584fd97dd94bcce74aa5fc052554ab7ace", + "gh_api_exit": 0, + "installed_equals_upstream": true + }, + { + "id": "or-aggressive", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/aggressive.ts", + "lines": [ + 123, + 145 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/aggressive.ts#L123-L145", + "claim": "Fallback summaries replace long messages.", + "installed_build": "dd6e9607e", + "evidence_class": "read-only installed source and gh-api byte comparison", + "live_acceptance": null, + "installed_file_sha256": "f9b3d0b6dcd9f441218b12663f706ff1cfe32b90d0f59aba291aa94b3adc7677", + "gh_api_exit": 0, + "installed_equals_upstream": true + }, + { + "id": "or-llmlingua", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/engines/llmlingua/index.ts", + "lines": [ + 417, + 489 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/llmlingua/index.ts#L417-L489", + "claim": "Native model compression path and runtime availability guard; enabled is not proof of application.", + "installed_build": "dd6e9607e", + "evidence_class": "read-only installed source and gh-api byte comparison", + "live_acceptance": null, + "installed_file_sha256": "f6ac8a7051d12bc3dd9abcd3a76df060d7855e78bb03367503c785549a40f11e", + "gh_api_exit": 0, + "installed_equals_upstream": true + }, + { + "id": "or-ultra", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/ultra.ts", + "lines": [ + 42, + 65 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/ultra.ts#L42-L65", + "claim": "Prose pruning protects selected blocks, not every lexical value.", + "installed_build": "dd6e9607e", + "evidence_class": "read-only installed source and gh-api byte comparison", + "live_acceptance": null, + "installed_file_sha256": "11484549092368163d277e250ad3a6bd99cfa38ba494aa8186492ba893468a0c", + "gh_api_exit": 0, + "installed_equals_upstream": true + }, + { + "id": "or-codex-transport", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/executors/codex.ts", + "lines": [ + 1498, + 1573 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/executors/codex.ts#L1498-L1573", + "claim": "Cache key injection, session body field removal, conditional encrypted reasoning policy and Responses allowlist. Also 365-391,1176-1218,1395-1405.", + "installed_build": "dd6e9607e", + "evidence_class": "read-only installed source and gh-api byte comparison", + "live_acceptance": null, + "installed_file_sha256": "8ca56325adc9775305464ea7adcfa9b4e70e18ce47825aa58f8c8929418f1289", + "gh_api_exit": 0, + "installed_equals_upstream": true + }, + { + "id": "or-call-logs", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "src/lib/usage/callLogs.ts", + "lines": [ + 90, + 107 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107", + "claim": "Source declares row id, duration, tokens_in/out/cache_read/reasoning; live database not inspected.", + "installed_build": "dd6e9607e", + "evidence_class": "read-only installed source and gh-api byte comparison", + "live_acceptance": null, + "installed_file_sha256": "64269db02cb0ed73f6f7d0a3daaf314a1aa08df2993accd559ec078aff764e9c", + "gh_api_exit": 0, + "installed_equals_upstream": true + }, + { + "id": "harbor-codex", + "repository": "harbor-framework/harbor", + "pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "file": "src/harbor/agents/installed/codex.py", + "lines": [ + 1207, + 1296 + ], + "url": "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1207-L1296", + "claim": "Native config loader preserves provider settings, merges runtime inputs, uploads config. CODEX_HOME 81,1352-1359; options 38-59; launch 1432-1448; native transcripts 1452-1476.", + "gh_api_exit": 0 + }, + { + "id": "harbor-options", + "repository": "harbor-framework/harbor", + "pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "file": "src/harbor/agents/options.py", + "lines": [ + 50, + 70 + ], + "url": "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/options.py#L50-L70", + "claim": "Profile forwarding not found in shared options, CodexOptions or native launch; arbitrary extras forbidden.", + "gh_api_exit": 0 + }, + { + "id": "harbor-verifier", + "repository": "harbor-framework/harbor", + "pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "file": "src/harbor/verifier/verifier.py", + "lines": [ + 165, + 250 + ], + "url": "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/verifier/verifier.py#L165-L250", + "claim": "Native task test execution and reward/artifact retention.", + "gh_api_exit": 0 + }, + { + "id": "rewardkit-example", + "repository": "harbor-framework/harbor", + "pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "file": "examples/tasks/reward-kit-example/tests/test.sh", + "lines": [ + 1, + 2 + ], + "url": "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/examples/tasks/reward-kit-example/tests/test.sh#L1-L2", + "claim": "Unchanged example pins harbor-rewardkit 0.1.8; bundled package version 0.2.0 is a different pin.", + "gh_api_exit": 0 + }, + { + "id": "promptfoo-responses", + "repository": "promptfoo/promptfoo", + "pin": "34f74d34e140b5e17d23770dfb2340057b1936b8", + "file": "src/providers/openai/responses.ts", + "lines": [ + 1195, + 1210 + ], + "url": "https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/providers/openai/responses.ts#L1195-L1210", + "claim": "POST /responses uses config.headers; request fields at 1003-1040.", + "gh_api_exit": 0 + }, + { + "id": "promptfoo-base-url", + "repository": "promptfoo/promptfoo", + "pin": "34f74d34e140b5e17d23770dfb2340057b1936b8", + "file": "src/providers/openai/index.ts", + "lines": [ + 111, + 128 + ], + "url": "https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/providers/openai/index.ts#L111-L128", + "claim": "apiBaseUrl selects custom OpenAI endpoint.", + "gh_api_exit": 0 + }, + { + "id": "promptfoo-cli", + "repository": "promptfoo/promptfoo", + "pin": "34f74d34e140b5e17d23770dfb2340057b1936b8", + "file": "src/commands/eval.ts", + "lines": [ + 54, + 97 + ], + "url": "https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/commands/eval.ts#L54-L97", + "claim": "--model-outputs supports offline grading; --repeat, --max-concurrency and --no-cache.", + "gh_api_exit": 0 + }, + { + "id": "promptfoo-python", + "repository": "promptfoo/promptfoo", + "pin": "34f74d34e140b5e17d23770dfb2340057b1936b8", + "file": "src/assertions/python.ts", + "lines": [ + 32, + 63 + ], + "url": "https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/assertions/python.ts#L32-L63", + "claim": "Python assertions normalize grading result; exceptions fail.", + "gh_api_exit": 0 + }, + { + "id": "promptfoo-startup", + "repository": "promptfoo/promptfoo", + "pin": "34f74d34e140b5e17d23770dfb2340057b1936b8", + "file": "src/main.ts", + "lines": [ + 63, + 64 + ], + "url": "https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/main.ts#L63-L64", + "claim": "Migration runs before argument parsing at 147; logger.ts 224-247 creates/cleans logs.", + "gh_api_exit": 0 + }, + { + "id": "inspect", + "repository": "UKGovernmentBEIS/inspect_ai", + "pin": "c2b63a0b0560f9c3b5b5230365e0a8fe08cd3df0", + "file": "CHANGELOG.md", + "lines": [ + 1, + 4 + ], + "url": "https://github.com/UKGovernmentBEIS/inspect_ai/blob/c2b63a0b0560f9c3b5b5230365e0a8fe08cd3df0/CHANGELOG.md#L1-L4", + "claim": "0.3.271 adds extra headers/body CLI; release endpoints 404, tagged changelog retrieved.", + "gh_api_exit": 0 + }, + { + "id": "inspect-swe", + "repository": "meridianlabs-ai/inspect_swe", + "pin": "7eb8dd64309db4cd0f6bdf1d0ffd9786a74a4088", + "file": "src/inspect_swe/_codex_cli/codex_cli.py", + "lines": [ + 619, + 637 + ], + "url": "https://github.com/meridianlabs-ai/inspect_swe/blob/7eb8dd64309db4cd0f6bdf1d0ffd9786a74a4088/src/inspect_swe/_codex_cli/codex_cli.py#L619-L637", + "claim": "0.2.71 rewrites provider and CODEX_HOME to Inspect bridge; cannot silently substitute transport.", + "gh_api_exit": 0 + }, + { + "id": "scipy", + "repository": "scipy/scipy", + "pin": "e4e854eaa8f18d807cd3496028e257e36caa93cc", + "file": "scipy/stats/_resampling.py", + "lines": [ + 300, + 394 + ], + "url": "https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_resampling.py#L300-L394", + "claim": "SciPy 1.18.1 bootstrap paired option; permutation_test at 1679-1780 and paired samples at 1892-1916.", + "gh_api_exit": 0 + }, + { + "id": "scipy-binomial", + "repository": "scipy/scipy", + "pin": "e4e854eaa8f18d807cd3496028e257e36caa93cc", + "file": "scipy/stats/_binomtest.py", + "lines": [ + 52, + 126 + ], + "url": "https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_binomtest.py#L52-L126", + "claim": "Exact proportion interval; binomtest at 187. Conservative harmful-discordance gate avoids zero-width bootstrap certification.", + "gh_api_exit": 0 + }, + { + "id": "statsmodels", + "repository": "statsmodels/statsmodels", + "pin": "278ff9950636cdd4939b4055e339a8e681d79cab", + "file": "statsmodels/stats/multitest.py", + "lines": [ + 99, + 149 + ], + "url": "https://github.com/statsmodels/statsmodels/blob/278ff9950636cdd4939b4055e339a8e681d79cab/statsmodels/stats/multitest.py#L99-L149", + "claim": "statsmodels 0.15.0 multipletests(method=holm), implementation 233-245.", + "gh_api_exit": 0 + }, + { + "id": "role-policy", + "evidence_class": "user instruction", + "repository": null, + "pin": null, + "file": null, + "lines": null, + "url": null, + "claim": "H6 is a normative role gate, not an OmniRoute capability." + }, + { + "id": "codex-provider", + "repository": "openai/codex", + "pin": "rust-v0.157.1", + "file": "codex-rs/model-provider-info/src/lib.rs", + "lines": [ + 137, + 173 + ], + "url": "https://github.com/openai/codex/blob/rust-v0.157.1/codex-rs/model-provider-info/src/lib.rs#L137-L173", + "gh_api_exit": 0, + "claim": "Native provider base_url, wire_api, http_headers and env_http_headers; header map at 385. Tagged release notes checked, installed client reports 0.157.1." + }, + { + "id": "numpy-rng", + "repository": "numpy/numpy", + "pin": "c5ab79c14c98bfda1e60770ffa23a6130f8267b7", + "file": "numpy/random/_pcg64.pyx", + "lines": [ + 53, + 108 + ], + "url": "https://github.com/numpy/numpy/blob/c5ab79c14c98bfda1e60770ffa23a6130f8267b7/numpy/random/_pcg64.pyx#L53-L108", + "gh_api_exit": 0, + "claim": "NumPy 2.4.0 PCG64 fixed-seed bitstream; freeze exact Generator distribution/runtime too. SciPy 1.18.1 pyproject accepts numpy>=2.0,<2.8." + }, + { + "id": "or-affinity-header", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "src/sse/handlers/chat.ts", + "lines": [ + 684, + 692 + ], + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/handlers/chat.ts#L684-L692", + "evidence_class": "read-only installed source review", + "claim": "request.headers.get(\"x-omniroute-connection\") is the requested connection pin; cross-hop identity mapping unobserved." + }, + { + "id": "harbor-oracle", + "repository": "harbor-framework/harbor", + "pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "file": "src/harbor/agents/oracle.py", + "lines": [ + 57, + 104 + ], + "url": "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/oracle.py#L57-L104", + "gh_api_exit": 0, + "claim": "Native reference solution agent; nop.py is native negative control." + }, + { + "id": "harbor-command", + "repository": "harbor-framework/harbor", + "pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "file": "src/harbor/agents/installed/codex.py", + "lines": [ + 1339, + 1449 + ], + "claim": "Last model segment; config upload; forced bypass, unified_exec, max option and resume.", + "url": "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1339-L1449", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "harbor-main-model", + "repository": "harbor-framework/harbor", + "pin": "3c82380859d187957cfd5cd64802b076d9779550", + "file": "src/harbor/agents/installed/codex.py", + "lines": [ + 1502, + 1605 + ], + "claim": "Inspected main still strips the prefix and passes the last segment to --model.", + "url": "https://github.com/harbor-framework/harbor/blob/3c82380859d187957cfd5cd64802b076d9779550/src/harbor/agents/installed/codex.py#L1502-L1605", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "codex-profile-merge", + "repository": "openai/codex", + "pin": "36650394c5b38c2990ccf2a3457165ca3e9d9726", + "file": "codex-rs/config/src/loader/mod.rs", + "lines": [ + 286, + 340 + ], + "claim": "Profile overlays base; legacy same-name profile selector/table rejected with -p.", + "url": "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/config/src/loader/mod.rs#L286-L340", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "codex-merge-semantics", + "repository": "openai/codex", + "pin": "36650394c5b38c2990ccf2a3457165ca3e9d9726", + "file": "codex-rs/config/src/merge.rs", + "lines": [ + 56, + 185 + ], + "claim": "Native deep merge includes key normalization and shell-filter replacement rules; plain dictionary overlay is insufficient.", + "url": "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/config/src/merge.rs#L56-L185", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "harbor-compose", + "repository": "harbor-framework/harbor", + "pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "file": "src/harbor/environments/docker/docker.py", + "lines": [ + 350, + 420 + ], + "claim": "Job extra_docker_compose and task-authored network_mode are respected.", + "url": "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/environments/docker/docker.py#L350-L420", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "harbor-compose-example", + "repository": "harbor-framework/harbor", + "pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "file": "examples/jobs/extra-docker-compose/config.yaml", + "lines": [ + 1, + 8 + ], + "claim": "Supported job-level overlay field.", + "url": "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/examples/jobs/extra-docker-compose/config.yaml#L1-L8", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "harbor-multi-step", + "repository": "harbor-framework/harbor", + "pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "file": "src/harbor/trial/multi_step.py", + "lines": [ + 25, + 110 + ], + "claim": "Native [[steps]] and resume_trajectory support checks.", + "url": "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/trial/multi_step.py#L25-L110", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "harbor-two-step-example", + "repository": "harbor-framework/harbor", + "pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "file": "examples/tasks/hello-multi-step-simple/task.toml", + "lines": [ + 15, + 31 + ], + "claim": "Exactly two original builder steps, create-file and append-content.", + "url": "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/examples/tasks/hello-multi-step-simple/task.toml#L15-L31", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "harbor-env", + "repository": "harbor-framework/harbor", + "pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "file": "src/harbor/agents/installed/base.py", + "lines": [ + 560, + 648 + ], + "claim": "extra_env and environment precedence; keep credential names in the configuration.", + "url": "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/base.py#L560-L648", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "harbor-redaction", + "repository": "harbor-framework/harbor", + "pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "file": "src/harbor/utils/env.py", + "lines": [ + 4, + 65 + ], + "claim": "KEY names are sensitive; use environment templates, never literal values or partial redaction.", + "url": "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/utils/env.py#L4-L65", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "promptfoo-codex-sdk", + "repository": "promptfoo/promptfoo", + "pin": "34f74d34e140b5e17d23770dfb2340057b1936b8", + "file": "src/providers/openai/codex-sdk.ts", + "lines": [ + 1034, + 1134 + ], + "claim": "config.model passed intact to thread options; explicit thread_id resumes, deep_tracing forces a fresh thread.", + "url": "https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/providers/openai/codex-sdk.ts#L1034-L1134", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "codex-sdk-model", + "repository": "openai/codex", + "pin": "36650394c5b38c2990ccf2a3457165ca3e9d9726", + "file": "sdk/typescript/src/exec.ts", + "lines": [ + 91, + 178 + ], + "claim": "SDK passes args.model intact after --model; configurable sandbox and config; explicit resume.", + "url": "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/sdk/typescript/src/exec.ts#L91-L178", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "inspect-swe-model", + "repository": "meridianlabs-ai/inspect_swe", + "pin": "7eb8dd64309db4cd0f6bdf1d0ffd9786a74a4088", + "file": "src/inspect_swe/_codex_cli/codex_cli.py", + "lines": [ + 483, + 637 + ], + "claim": "--model uses resolved Codex catalog slug; provider points at Inspect bridge, not directly at the lane.", + "url": "https://github.com/meridianlabs-ai/inspect_swe/blob/7eb8dd64309db4cd0f6bdf1d0ffd9786a74a4088/src/inspect_swe/_codex_cli/codex_cli.py#L483-L637", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "or-correlation", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "src/sse/handlers/chatHelpers.ts", + "lines": [ + 1172, + 1184 + ], + "claim": "Entry response returns X-Correlation-Id.", + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/handlers/chatHelpers.ts#L1172-L1184", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "or-effort", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "src/lib/usage/callLogs.ts", + "lines": [ + 628, + 653 + ], + "claim": "Effort columns are conditional on encrypted reasoning; body reasoning.effort is authoritative.", + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L628-L653", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "or-detail-api", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "src/app/api/usage/call-logs/[id]/route.ts", + "lines": [ + 1, + 22 + ], + "claim": "Read-only GET for detail by log id; response schema must be qualified on host.", + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/app/api/usage/call-logs/[id]/route.ts#L1-L22", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "or-affinity", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "src/sse/services/sessionAffinityPin.ts", + "lines": [ + 197, + 221 + ], + "claim": "Session headers, body session fields, prompt_cache_key, then input hash; connection pin is separate.", + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/services/sessionAffinityPin.ts#L197-L221", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "or-session-fallback", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/sessionManager.ts", + "lines": [ + 103, + 145 + ], + "claim": "Body/model/provider/first-user/tools/connection fingerprint fallback for live zone.", + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/sessionManager.ts#L103-L145", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "or-styles-catalog", + "repository": "diegosouzapw/OmniRoute", + "pin": "a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "file": "open-sse/services/compression/outputStyles/catalog.ts", + "lines": [ + 32, + 194 + ], + "claim": "Styles affect prose and code; JSON tasks do not support unrestricted prose-quality claims.", + "url": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/outputStyles/catalog.ts#L32-L194", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "codex-tool-shape", + "repository": "openai/codex", + "pin": "36650394c5b38c2990ccf2a3457165ca3e9d9726", + "file": "codex-rs/core/src/tools/context.rs", + "lines": [ + 524, + 601 + ], + "claim": "Unified exec prepends headers; function/custom output conversion is shape dependent.", + "url": "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/tools/context.rs#L524-L601", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "codex-shell-shape", + "repository": "openai/codex", + "pin": "36650394c5b38c2990ccf2a3457165ca3e9d9726", + "file": "codex-rs/core/src/tools/mod.rs", + "lines": [ + 97, + 124 + ], + "claim": "Legacy shell also wraps output with exit/time/header text.", + "url": "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/tools/mod.rs#L97-L124", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "codex-response-headers", + "repository": "openai/codex", + "pin": "36650394c5b38c2990ccf2a3457165ca3e9d9726", + "file": "codex-rs/codex-api/src/sse/responses.rs", + "lines": [ + 30, + 101 + ], + "claim": "Native reader exposes x-request-id and selected fields, not an observed X-Correlation-Id exporter.", + "url": "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/codex-api/src/sse/responses.rs#L30-L101", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "mitmproxy-headers", + "repository": "mitmproxy/mitmproxy", + "pin": "6c09d56e4c29a92f5ad01b03199977584b8ea14f", + "file": "mitmproxy/proxy/layers/http/_hooks.py", + "lines": [ + 7, + 39 + ], + "claim": "Header hooks run before bodies; supports a narrow response-header receipt adapter.", + "url": "https://github.com/mitmproxy/mitmproxy/blob/6c09d56e4c29a92f5ad01b03199977584b8ea14f/mitmproxy/proxy/layers/http/_hooks.py#L7-L39", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "mitmproxy-stream", + "repository": "mitmproxy/mitmproxy", + "pin": "6c09d56e4c29a92f5ad01b03199977584b8ea14f", + "file": "examples/addons/http-stream-simple.py", + "lines": [ + 1, + 14 + ], + "claim": "Unchanged upstream example enables response streaming.", + "url": "https://github.com/mitmproxy/mitmproxy/blob/6c09d56e4c29a92f5ad01b03199977584b8ea14f/examples/addons/http-stream-simple.py#L1-L14", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "mitmproxy-reverse", + "repository": "mitmproxy/mitmproxy", + "pin": "6c09d56e4c29a92f5ad01b03199977584b8ea14f", + "file": "docs/src/content/concepts/modes.md", + "lines": [ + 178, + 257 + ], + "claim": "Supported reverse mode forwards to a fixed entry gateway; no model rewrite needed.", + "url": "https://github.com/mitmproxy/mitmproxy/blob/6c09d56e4c29a92f5ad01b03199977584b8ea14f/docs/src/content/concepts/modes.md#L178-L257", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "scipy-exact", + "repository": "scipy/scipy", + "pin": "e4e854eaa8f18d807cd3496028e257e36caa93cc", + "file": "scipy/stats/_binomtest.py", + "lines": [ + 52, + 115 + ], + "claim": "Clopper-Pearson exact bounds supply a conservative paired net-difference fallback.", + "url": "https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_binomtest.py#L52-L115", + "evidence_class": "pinned source inspection; not runtime acceptance" + }, + { + "id": "tango", + "url": "https://doi.org/10.1002/(SICI)1097-0258(19980430)17:8%3C891::AID-SIM780%3E3.0.CO;2-B", + "claim": "Tango (1998), Equivalence test and confidence interval for the difference in proportions for the paired-sample design", + "evidence_class": "published reference; no new implementation or runtime result" + }, + { + "id": "intersection-union", + "url": "https://doi.org/10.1214/ss/1032280304", + "claim": "Berger and Hsu (1996), Bioequivalence trials, intersection-union tests and equivalence confidence sets", + "evidence_class": "published reference; no new implementation or runtime result" + }, + { + "id": "numpy-pilot-simulation", + "repository": "numpy/numpy", + "pin": "c5ab79c14c98bfda1e60770ffa23a6130f8267b7", + "version": "2.4.0", + "file": "numpy/random/_generator.pyx", + "lines": [ + 299, + 310 + ], + "url": "https://github.com/numpy/numpy/blob/c5ab79c14c98bfda1e60770ffa23a6130f8267b7/numpy/random/_generator.pyx#L299-L310", + "claim": "Generator.random uniform draws; integers at 576-587; multinomial at 3952-3963. These supply simulated law primitives, not a passed statistical pilot.", + "evidence_class": "pinned upstream gh-api source inspection" + } + ], + "client": { + "version": "0.157.1", + "profile": "stack-worker", + "effort": "max", + "argv_contract": [ + "codex", + "exec", + "[resume --last on turns 2/3]", + "--dangerously-bypass-approvals-and-sandbox", + "--skip-git-repo-check", + "--model", + "gpt-6-astra-max", + "--json", + "--enable", + "unified_exec", + "-c", + "model_reasoning_effort=max", + "--", + "" + ], + "stdin": "DEVNULL", + "same_config_except_provider": true, + "config_rule": "Materialize the native semantic deep merge of frozen sanitized base config.toml and stack-worker.config.toml. Pass that merged file through Harbor config. Hash both inputs, merged file, uploaded config, resolved effective configuration and binary. No auth-store copy; task/project/system layers and Harbor MCP/runtime overrides must match the reference.", + "static_provider_header": { + "A0/A1": { + "x-omniroute-compression": "allow-lossy" + }, + "D0/D1": {}, + "C": {} + }, + "wire_api": "responses", + "unsuffixed_allowed": false, + "unsuffixed_unlock": "Only a pre-confirmatory amendment with actual request-body reasoning.effort=max and route/metadata equivalence at both hops may change the suffixed target. Null effort columns never establish drift.", + "runtime_acceptance": null, + "provider_example": { + "model_provider": "ab_gateway", + "model_providers": { + "ab_gateway": { + "name": "Compression AB gateway", + "base_url": "", + "wire_api": "responses", + "requires_openai_auth": false, + "http_headers": { + "x-omniroute-compression": "" + }, + "env_key": "OMNIROUTE_FW_API_KEY" + } + }, + "source_id": "codex-provider", + "note": "Engines-on template only; C omits env_key. base_url is entry gateway, or the qualified client-side capture listener forwarding unchanged to it. Static http_headers hold compression switches only." + }, + "model_hops": { + "client_to_20129": "sharedgw/gpt-6-astra-max", + "20129_to_20128": "gpt-6-astra-max", + "control_C": "cx/gpt-6-astra-max" + }, + "shell_contract": "Harbor optionally sources ~/.nvm/nvm.sh; command redirects 2>&1 --model at L483-547; custom provider/bridge at L618-637. Native Codex executes but bridge translation needs full profile/transport/correlation qualification; not an intact direct route.", + "source_ids": [ + "inspect-swe-model" + ] + }, + { + "name": "promptfoo Codex SDK 0.123.1", + "finding": "config.model -> threadOptions.model unchanged at L1054-1058; openai SDK args.model -> --model at L113-115. Strong alternate for prefix preservation. Thread-id continuation exists; deep_tracing creates fresh threads. Native task verifier/container/three-turn and SDK pin equivalence remain unqualified.", + "source_ids": [ + "promptfoo-codex-sdk", + "codex-sdk-model" + ] + } + ], + "rationale": "Retain the source-backed native Harbor task/verifier/multi-step path under the user's option (b), with an explicit unresolved route gate. This is a conditional draft choice, not evidence Harbor currently works for these cells.", + "overturn_condition": "Before confirmatory data, compare promptfoo Codex SDK against Harbor on identical merged config, prompt-input, exact three-turn tools/verifiers, header joins and both-hop effort/metadata. Switch by amendment if the SDK passes these and Harbor has no supported slashless route, or has unequal native semantics. Never pool runners.", + "network": { + "source_ids": [ + "harbor-compose", + "harbor-compose-example" + ], + "job_field": "environment.extra_docker_compose", + "overlay": { + "services": { + "main": { + "network_mode": "host" + } + } + }, + "overlay_sha256": null, + "result": null, + "reachability_probe": "From the actual main container run a bounded TCP connection probe to 127.0.0.1:20128 and :20129, then an unauthenticated non-inference HTTP health probe with body discarded; retain only address class/status/latency. A TCP success alone is not gateway/provider acceptance. Qualify the WSL2/Docker host network implementation; do not assume it shares the host loopback.", + "admin_api_exposure": "Host networking plus --dangerously-bypass-approvals-and-sandbox exposes passwordless gateway management APIs and other reachable host services. A lane inference API key does not isolate management APIs.", + "bounds": "Use only frozen trusted synthetic tasks; no hostile checkout/internet task input, host directories, Docker socket or host credential mounts. One owned session, independent read-only effective-config snapshots before/after every session and immediate invalidation on drift. These are detection/scope bounds, not a security boundary. Live execution additionally requires host-owner qualification of an upstream-supported network allowlist that blocks management paths/direct alternate routes; no such containment is accepted here. If unavailable, do not execute this host-network design; amend to an isolated host/gateway setup before data." + } + }, + "gateway": { + "name": "promptfoo", + "version": "0.123.1", + "source_ids": [ + "promptfoo-responses", + "promptfoo-base-url", + "promptfoo-cli", + "promptfoo-python" + ], + "provider_id_template": "openai:responses:", + "config_fields": [ + "apiBaseUrl", + "headers", + "reasoning.effort", + "store", + "include", + "prompt_cache_key" + ], + "command_contract": "promptfoo eval --no-cache --repeat --max-concurrency 1; offline --model-outputs for native Codex answers", + "boundary": "Direct promptfoo requests are gateway diagnostics, not native Codex role attempts or proof of multi-turn tool use.", + "model_examples": [ + "openai:responses:cx/gpt-6-astra-max", + "openai:responses:sharedgw/gpt-6-astra-max" + ] + }, + "runner_up": { + "name": "Inspect 0.3.271 + inspect_swe 0.2.71", + "source_ids": [ + "inspect", + "inspect-swe" + ], + "readiness": "See primary.alternatives. Inspect bridge is not interchangeable with the native direct-gateway path; source and host qualification required before any amendment." + }, + "rewardkit": { + "bundled_version": "0.2.0", + "unchanged_example_dependency": "0.1.8", + "source_id": "rewardkit-example", + "readiness": "Freeze actual criteria and dependency; do not relabel 0.1.8 example as 0.2.0 execution." + }, + "statistics": { + "scipy": "1.18.1", + "statsmodels": "0.15.0", + "installed_acceptance": null, + "numpy": "2.4.0" + } + }, + "cells": [ + { + "id": "C", + "classification": "confirmatory", + "base_url": "http://127.0.0.1:20128/v1", + "model": "cx/gpt-6-astra-max", + "engines": [], + "compression_header": null, + "output_styles_on": false, + "principal": "keyless", + "live_zone_configured": false, + "observed_effective_plan": null, + "canonical_model_is_runner_argument": false + }, + { + "id": "D0", + "classification": "confirmatory", + "base_url": "http://127.0.0.1:20129/v1", + "model": "sharedgw/gpt-6-astra-max", + "engines": [ + "session-dedup", + "ccr", + "lite", + "headroom" + ], + "compression_header": null, + "output_styles_on": false, + "principal": "registered-lane-key", + "live_zone_configured": true, + "observed_effective_plan": null, + "canonical_model_is_runner_argument": false, + "upstream_model": "gpt-6-astra-max", + "env_key": "OMNIROUTE_FW_API_KEY", + "live_zone_ttl_minutes": 60 + }, + { + "id": "D1", + "classification": "confirmatory", + "base_url": "http://127.0.0.1:20129/v1", + "model": "sharedgw/gpt-6-astra-max", + "engines": [ + "session-dedup", + "ccr", + "lite", + "headroom" + ], + "compression_header": null, + "output_styles_on": true, + "principal": "registered-lane-key", + "live_zone_configured": true, + "observed_effective_plan": null, + "canonical_model_is_runner_argument": false, + "upstream_model": "gpt-6-astra-max", + "env_key": "OMNIROUTE_FW_API_KEY", + "live_zone_ttl_minutes": 60 + }, + { + "id": "A0", + "classification": "confirmatory", + "base_url": "http://127.0.0.1:20129/v1", + "model": "sharedgw/gpt-6-astra-max", + "engines": [ + "session-dedup", + "ccr", + "lite", + "rtk", + "codex-responses", + "headroom", + "relevance", + "caveman", + "aggressive", + "llmlingua", + "ultra", + "omniglyph" + ], + "compression_header": "allow-lossy", + "output_styles_on": false, + "principal": "registered-lane-key", + "live_zone_configured": true, + "observed_effective_plan": null, + "canonical_model_is_runner_argument": false, + "upstream_model": "gpt-6-astra-max", + "env_key": "OMNIROUTE_FW_API_KEY", + "live_zone_ttl_minutes": 60 + }, + { + "id": "A1", + "classification": "confirmatory", + "base_url": "http://127.0.0.1:20129/v1", + "model": "sharedgw/gpt-6-astra-max", + "engines": [ + "session-dedup", + "ccr", + "lite", + "rtk", + "codex-responses", + "headroom", + "relevance", + "caveman", + "aggressive", + "llmlingua", + "ultra", + "omniglyph" + ], + "compression_header": "allow-lossy", + "output_styles_on": true, + "principal": "registered-lane-key", + "live_zone_configured": true, + "observed_effective_plan": null, + "canonical_model_is_runner_argument": false, + "upstream_model": "gpt-6-astra-max", + "env_key": "OMNIROUTE_FW_API_KEY", + "live_zone_ttl_minutes": 60 + } + ], + "factorization": { + "design": "Direct clean control plus 2x2 crossing upstream safe-labelled default/all engines with output styles off/on. Five cells; four one-factor lossy/style contrasts avoid confounding an output change with input compression.", + "header_off_is_clean_control": false, + "direct_control": { + "compression_enabled": false, + "exclusions": [ + "codex/*" + ] + }, + "framework_node": "sharedgw -> 20128; client canonical one-slash slug, upstream bare gpt-6-astra-max; source-supported runner route still gated.", + "styles_on": [ + { + "id": "terse-prose", + "level": "full" + }, + { + "id": "less-code", + "level": "full" + }, + { + "id": "ponytail", + "level": "full" + }, + { + "id": "i-have-adhd", + "level": "full" + } + ], + "legacy_on": { + "enabled": true, + "intensity": "lite" + }, + "styles_off": { + "outputStyles": [], + "cavemanOutputMode": { + "enabled": false + } + }, + "style_resolution": "Nonempty explicit styles win; legacy lite is configured but not added as a fifth style.", + "allocation": "All four comparisons to C in each role are confirmatory. Internal engine/style contrasts and interaction estimates are explanatory, without promotion claims.", + "configuration_application": "Future host owner snapshots and read-backs 20129 settings for each cell in exclusive experiment windows, preserving all unrelated fields; restore exact snapshot after block. No live changes while drafting. If exclusive scope unavailable, do not run mixed-style cells on the shared instance.", + "engine_activation": "Pin effective config and model dependencies, verify applied vs no-eligible/skipped per engine. All-enabled treatment is not evidence that every engine executed.", + "sharedgw_wire_api": "responses", + "wire_qualification": "Freeze provider-node source_format/target_format and actual Responses -> Responses wire at both hops. No Chat fallback accepted for confirmation; verify encrypted reasoning and custom tool preservation. A fallback requires amendment/new qualification.", + "keyed_live_zone": { + "cells": [ + "D0", + "D1", + "A0", + "A1" + ], + "classification": "confirmatory", + "ttl_minutes": 60, + "settings_field": "cacheMinutes", + "key_file_mode": "0600", + "env_key": "OMNIROUTE_FW_API_KEY", + "source_ids": [ + "or-live-zone", + "or-chat-core", + "harbor-env" + ], + "configuration_fact": "User-established existing registered lane key and TTL 60 minutes; not queried, created or opened by repair.", + "principal_receipt": "Only salted principal ID and equality booleans; registered configuration is known, per-session live-zone reuse remains unobserved.", + "live_result": null + } + }, + "exploratory": { + "allows_promotion": false, + "single_engine_order": [ + "codex-responses", + "rtk", + "relevance", + "caveman", + "aggressive", + "llmlingua", + "ultra", + "omniglyph" + ], + "single_engine_design": "If remaining reserved budget supports a complete paired block, add exactly one listed engine to D0, styles off; compare D0 within same block, 3 repetitions x 6 tasks x 3 roles. Fixed order, no data-driven engine selection. Never infer all-engine interactions from individual passes. New confirmation required before adopting any such plan.", + "budget_rule": "No live-zone exploration: keyed is confirmatory. Optional single-engine blocks use a separately pilot-sized allocation after complete confirmation; zero allocation is permitted, no promotion.", + "single_engine_repetition_unit": "3 persistent three-turn sessions per task x 6 tasks x 3 roles per added-engine cell, each paired with D0; complete blocks only." + }, + "hashing": { + "packet": "UTF-8 JSON, sort_keys=true, ensure_ascii=true, separators=(comma,colon), allow_nan=false, no trailing newline; hash packet object only.", + "upstream_task_directory": "All files sorted bytewise by relative path; concatenate relative_path + NUL + sha256(file_bytes).hexdigest() + LF; sha256(concatenation).", + "freeze_scope": "Draft task selections and literal stimuli have hashes now. Runtime images/dependencies/adapter executables and final execution seal remain null; no provider request before sealing." + }, + "task_sets": { + "builder": { + "packet": { + "role": "builder", + "tasks": [ + { + "id": "B-deepswe_tomlkit-toml-table-converters", + "prompt": "Implement TOML table converters; preserve values/comments and specified error behavior.", + "prompt_binding": "Original instruction/step and verifier bytes retained; additional canary instructions and three-step envelope are explicitly local integration. See multi_turn.task_turn_mapping; do not claim unchanged upstream execution for combined trial.", + "source_ids": [ + "harbor-codex", + "harbor-verifier" + ], + "source_repository": "harbor-framework/harbor", + "source_pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "source_directory": "examples/tasks/deepswe_tomlkit-toml-table-converters", + "source_files": 8, + "source_manifest_sha256": "fa8dbedf2d1a53a7795b8ef4978742e87eaf20f5e50ff822574d9477fcdfae5e", + "oracle": "Unchanged task tests/test.sh or each native step verifier; full reward contract from pinned task, all required checks pass, no skipped/missing/nonfinite reward.", + "task_bytes_verified_by": "gh api: all source blobs read in memory; source-byte manifest hash independently recomputed", + "runtime_image_digest": null + }, + { + "id": "B-reward-kit-example", + "prompt": "Implement textstats and JSON analysis pipeline against unchanged Rewardkit criteria.", + "prompt_binding": "Original instruction/step and verifier bytes retained; additional canary instructions and three-step envelope are explicitly local integration. See multi_turn.task_turn_mapping; do not claim unchanged upstream execution for combined trial.", + "source_ids": [ + "harbor-codex", + "harbor-verifier", + "rewardkit-example" + ], + "source_repository": "harbor-framework/harbor", + "source_pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "source_directory": "examples/tasks/reward-kit-example", + "source_files": 15, + "source_manifest_sha256": "28959a602e7c90928cae7535ce90b48f4a8d2a67bddabc0dc190cc772e7c2661", + "oracle": "Unchanged task tests/test.sh or each native step verifier; full reward contract from pinned task, all required checks pass, no skipped/missing/nonfinite reward.", + "task_bytes_verified_by": "gh api: all source blobs read in memory; source-byte manifest hash independently recomputed", + "runtime_image_digest": null + }, + { + "id": "B-hello-multi-step-simple", + "prompt": "Execute native create-file and append-content steps, preserving each native verifier.", + "prompt_binding": "Original instruction/step and verifier bytes retained; additional canary instructions and three-step envelope are explicitly local integration. See multi_turn.task_turn_mapping; do not claim unchanged upstream execution for combined trial.", + "source_ids": [ + "harbor-codex", + "harbor-verifier" + ], + "source_repository": "harbor-framework/harbor", + "source_pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "source_directory": "examples/tasks/hello-multi-step-simple", + "source_files": 8, + "source_manifest_sha256": "5da0ec364eb13199c924e9daf7826044276ad453e04aca7fb07a8d2b19476a0e", + "oracle": "Unchanged task tests/test.sh or each native step verifier; full reward contract from pinned task, all required checks pass, no skipped/missing/nonfinite reward.", + "task_bytes_verified_by": "gh api: all source blobs read in memory; source-byte manifest hash independently recomputed", + "runtime_image_digest": null + }, + { + "id": "B-hello-mcp", + "prompt": "Use task MCP get_secret and persist the exact returned synthetic value.", + "prompt_binding": "Original instruction/step and verifier bytes retained; additional canary instructions and three-step envelope are explicitly local integration. See multi_turn.task_turn_mapping; do not claim unchanged upstream execution for combined trial.", + "source_ids": [ + "harbor-codex", + "harbor-verifier" + ], + "source_repository": "harbor-framework/harbor", + "source_pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "source_directory": "examples/tasks/hello-mcp", + "source_files": 10, + "source_manifest_sha256": "f420d962ff4c12ffd6cff758f1f81c26c7f8c4361cf14d7310cf09441a549565", + "oracle": "Unchanged task tests/test.sh or each native step verifier; full reward contract from pinned task, all required checks pass, no skipped/missing/nonfinite reward.", + "task_bytes_verified_by": "gh api: all source blobs read in memory; source-byte manifest hash independently recomputed", + "runtime_image_digest": null + }, + { + "id": "B-hello-workdir", + "prompt": "Write the working directory as required by the original instruction.", + "prompt_binding": "Original instruction/step and verifier bytes retained; additional canary instructions and three-step envelope are explicitly local integration. See multi_turn.task_turn_mapping; do not claim unchanged upstream execution for combined trial.", + "source_ids": [ + "harbor-codex", + "harbor-verifier" + ], + "source_repository": "harbor-framework/harbor", + "source_pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "source_directory": "examples/tasks/hello-workdir", + "source_files": 5, + "source_manifest_sha256": "f8fddbfe39d4c1375391155cd7a698b7ae6af2637b8738f0283848bd2061db6d", + "oracle": "Unchanged task tests/test.sh or each native step verifier; full reward contract from pinned task, all required checks pass, no skipped/missing/nonfinite reward.", + "task_bytes_verified_by": "gh api: all source blobs read in memory; source-byte manifest hash independently recomputed", + "runtime_image_digest": null + }, + { + "id": "B-separate-verifier-environment", + "prompt": "Produce specified artifacts for the original isolated verifier.", + "prompt_binding": "Original instruction/step and verifier bytes retained; additional canary instructions and three-step envelope are explicitly local integration. See multi_turn.task_turn_mapping; do not claim unchanged upstream execution for combined trial.", + "source_ids": [ + "harbor-codex", + "harbor-verifier" + ], + "source_repository": "harbor-framework/harbor", + "source_pin": "1e5c5c6db929a10a140d05e606882c671ae20729", + "source_directory": "examples/tasks/separate-verifier-environment", + "source_files": 6, + "source_manifest_sha256": "4a0001756486d276c78e01aefb3ee901617ccb99e8ca2c54acf29d99f50eec0a", + "oracle": "Unchanged task tests/test.sh or each native step verifier; full reward contract from pinned task, all required checks pass, no skipped/missing/nonfinite reward.", + "task_bytes_verified_by": "gh api: all source blobs read in memory; source-byte manifest hash independently recomputed", + "runtime_image_digest": null + } + ], + "controls": [ + { + "kind": "known-pass", + "input": { + "native_agent": "oracle", + "task_set": "all six unchanged tasks" + }, + "expected_accept": true + }, + { + "kind": "known-fail", + "input": { + "native_agent": "nop", + "task_set": "all six unchanged tasks" + }, + "expected_accept": false + }, + { + "kind": "malformed", + "input": { + "reward_artifact_bytes": "{\"reward\": NaN}", + "variant": "missing/truncated verifier artifact also required" + }, + "expected_accept": false + } + ], + "canary_packet_sha256": "f06487b5af0fdab5a12dbb35cae00b4b6b0e6ec938311a961c9f5ea332446b43", + "shared_files": {}, + "delivery": "Native Codex task + same-session three-user-turn tool canary; promptfoo grades retained outputs offline. Upstream task verifiers retain their own separate rewards.", + "scope": "Frozen finite task set only; no population-wide or arbitrary repository role claim." + }, + "sha256": "4ddbb5aae74a1fa6b613d836f62120abf28f3ffd1091ae90e496edb9b879444b", + "materialized_runtime_sha256": null, + "native_controls_result": null + }, + "reviewer": { + "packet": { + "role": "reviewer", + "tasks": [ + { + "id": "R-number", + "prompt": "Read the supplied patch file and companion tool evidence. Find the one planted defect under the frozen contract. Return only {\"defect_id\": string, \"file\": string, \"exact_values\": string[]}; do not edit files. Explain no additional speculative defects. Allowed defect_id vocabulary: JSON_NUMBER_ROUNDING, MESSAGE_KEY_COLLISION, TOOL_TAIL_LOSS, CACHE_DOUBLE_COUNT, LEGACY_STYLE_STILL_ENABLED, ENCRYPTED_REASONING_DROPPED, NONE. Set file to the supplied patch filename. exact_values must contain, in order: original integer lexeme, then original decimal lexeme from numeric.json. Obtain companion facts through tools; do not guess missing values.", + "source_ids": [ + "or-minify", + "promptfoo-python" + ], + "evidence_class": "synthetic mutation derived from cited upstream behavior", + "files": { + "numeric.ts": "export const compact = (text: string) => JSON.stringify(JSON.parse(text));\n" + }, + "companion_tool_input": { + "canary_packet_sha256": "8bf066389d8523cbfa1d3593d399a986083b6e875e652901f266469dce5a6d07", + "usage_row": { + "tokens_in": 1000, + "tokens_cache_read": 940 + } + }, + "expected": { + "defect_id": "JSON_NUMBER_ROUNDING", + "file": "numeric.ts", + "exact_values": [ + "1234567890123456711", + "1.50" + ] + }, + "oracle": "promptfoo native Python assertion: strict JSON parse, reject duplicate keys/nonfinite values, require exactly expected keys/types and byte-exact string values; exact object equality; malformed or missing output fails. Artifact-derived source facts, never model self-score.", + "schema": { + "type": "object", + "additionalProperties": false, + "required": [ + "defect_id", + "file", + "exact_values" + ], + "properties": { + "defect_id": { + "type": "string", + "enum": [ + "JSON_NUMBER_ROUNDING", + "MESSAGE_KEY_COLLISION", + "TOOL_TAIL_LOSS", + "CACHE_DOUBLE_COUNT", + "LEGACY_STYLE_STILL_ENABLED", + "ENCRYPTED_REASONING_DROPPED", + "NONE" + ] + }, + "file": { + "type": "string" + }, + "exact_values": { + "type": "array", + "items": { + "type": "string" + } + } + } + }, + "grader_controls": [ + { + "kind": "known-pass", + "output": "{\"defect_id\":\"JSON_NUMBER_ROUNDING\",\"file\":\"numeric.ts\",\"exact_values\":[\"1234567890123456711\",\"1.50\"]}", + "expected_accept": true + }, + { + "kind": "known-fail", + "output": "{\"defect_id\":\"WRONG_CONTROL_VALUE\",\"file\":\"numeric.ts\",\"exact_values\":[\"1234567890123456711\",\"1.50\"]}", + "expected_accept": false + }, + { + "kind": "malformed", + "output": "{\"unterminated\":", + "expected_accept": false + } + ] + }, + { + "id": "R-key", + "prompt": "Read the supplied patch file and companion tool evidence. Find the one planted defect under the frozen contract. Return only {\"defect_id\": string, \"file\": string, \"exact_values\": string[]}; do not edit files. Explain no additional speculative defects. Allowed defect_id vocabulary: JSON_NUMBER_ROUNDING, MESSAGE_KEY_COLLISION, TOOL_TAIL_LOSS, CACHE_DOUBLE_COUNT, LEGACY_STYLE_STILL_ENABLED, ENCRYPTED_REASONING_DROPPED, NONE. Set file to the supplied patch filename. exact_values must contain, in order: stringKey(2) formatted as string(2)=, then partKey(0,1) formatted as part(0,1)=. Obtain companion facts through tools; do not guess missing values.", + "source_ids": [ + "or-dedup", + "promptfoo-python" + ], + "evidence_class": "synthetic mutation derived from cited upstream behavior", + "files": { + "keys.ts": "const stringKey = (i: number) => i;\nconst partKey = (i: number, p: number) => i * 100000 + p + 1;\n" + }, + "companion_tool_input": { + "canary_packet_sha256": "8bf066389d8523cbfa1d3593d399a986083b6e875e652901f266469dce5a6d07", + "usage_row": { + "tokens_in": 1000, + "tokens_cache_read": 940 + } + }, + "expected": { + "defect_id": "MESSAGE_KEY_COLLISION", + "file": "keys.ts", + "exact_values": [ + "string(2)=2", + "part(0,1)=2" + ] + }, + "oracle": "promptfoo native Python assertion: strict JSON parse, reject duplicate keys/nonfinite values, require exactly expected keys/types and byte-exact string values; exact object equality; malformed or missing output fails. Artifact-derived source facts, never model self-score.", + "schema": { + "type": "object", + "additionalProperties": false, + "required": [ + "defect_id", + "file", + "exact_values" + ], + "properties": { + "defect_id": { + "type": "string", + "enum": [ + "JSON_NUMBER_ROUNDING", + "MESSAGE_KEY_COLLISION", + "TOOL_TAIL_LOSS", + "CACHE_DOUBLE_COUNT", + "LEGACY_STYLE_STILL_ENABLED", + "ENCRYPTED_REASONING_DROPPED", + "NONE" + ] + }, + "file": { + "type": "string" + }, + "exact_values": { + "type": "array", + "items": { + "type": "string" + } + } + } + }, + "grader_controls": [ + { + "kind": "known-pass", + "output": "{\"defect_id\":\"MESSAGE_KEY_COLLISION\",\"file\":\"keys.ts\",\"exact_values\":[\"string(2)=2\",\"part(0,1)=2\"]}", + "expected_accept": true + }, + { + "kind": "known-fail", + "output": "{\"defect_id\":\"WRONG_CONTROL_VALUE\",\"file\":\"keys.ts\",\"exact_values\":[\"string(2)=2\",\"part(0,1)=2\"]}", + "expected_accept": false + }, + { + "kind": "malformed", + "output": "{\"unterminated\":", + "expected_accept": false + } + ] + }, + { + "id": "R-tail", + "prompt": "Read the supplied patch file and companion tool evidence. Find the one planted defect under the frozen contract. Return only {\"defect_id\": string, \"file\": string, \"exact_values\": string[]}; do not edit files. Explain no additional speculative defects. Allowed defect_id vocabulary: JSON_NUMBER_ROUNDING, MESSAGE_KEY_COLLISION, TOOL_TAIL_LOSS, CACHE_DOUBLE_COUNT, LEGACY_STYLE_STILL_ENABLED, ENCRYPTED_REASONING_DROPPED, NONE. Set file to the supplied patch filename. exact_values must contain, in order: the cutoff as a decimal integer string, then TAIL_PIN from long.txt. Obtain companion facts through tools; do not guess missing values.", + "source_ids": [ + "or-lite", + "promptfoo-python" + ], + "evidence_class": "synthetic mutation derived from cited upstream behavior", + "files": { + "tool.ts": "const shrink = (text: string) => text.length > 2000 ? text.slice(0, 2000) : text;\n" + }, + "companion_tool_input": { + "canary_packet_sha256": "8bf066389d8523cbfa1d3593d399a986083b6e875e652901f266469dce5a6d07", + "usage_row": { + "tokens_in": 1000, + "tokens_cache_read": 940 + } + }, + "expected": { + "defect_id": "TOOL_TAIL_LOSS", + "file": "tool.ts", + "exact_values": [ + "2000", + "dd6e9607e" + ] + }, + "oracle": "promptfoo native Python assertion: strict JSON parse, reject duplicate keys/nonfinite values, require exactly expected keys/types and byte-exact string values; exact object equality; malformed or missing output fails. Artifact-derived source facts, never model self-score.", + "schema": { + "type": "object", + "additionalProperties": false, + "required": [ + "defect_id", + "file", + "exact_values" + ], + "properties": { + "defect_id": { + "type": "string", + "enum": [ + "JSON_NUMBER_ROUNDING", + "MESSAGE_KEY_COLLISION", + "TOOL_TAIL_LOSS", + "CACHE_DOUBLE_COUNT", + "LEGACY_STYLE_STILL_ENABLED", + "ENCRYPTED_REASONING_DROPPED", + "NONE" + ] + }, + "file": { + "type": "string" + }, + "exact_values": { + "type": "array", + "items": { + "type": "string" + } + } + } + }, + "grader_controls": [ + { + "kind": "known-pass", + "output": "{\"defect_id\":\"TOOL_TAIL_LOSS\",\"file\":\"tool.ts\",\"exact_values\":[\"2000\",\"dd6e9607e\"]}", + "expected_accept": true + }, + { + "kind": "known-fail", + "output": "{\"defect_id\":\"WRONG_CONTROL_VALUE\",\"file\":\"tool.ts\",\"exact_values\":[\"2000\",\"dd6e9607e\"]}", + "expected_accept": false + }, + { + "kind": "malformed", + "output": "{\"unterminated\":", + "expected_accept": false + } + ] + }, + { + "id": "R-cache", + "prompt": "Read the supplied patch file and companion tool evidence. Find the one planted defect under the frozen contract. Return only {\"defect_id\": string, \"file\": string, \"exact_values\": string[]}; do not edit files. Explain no additional speculative defects. Allowed defect_id vocabulary: JSON_NUMBER_ROUNDING, MESSAGE_KEY_COLLISION, TOOL_TAIL_LOSS, CACHE_DOUBLE_COUNT, LEGACY_STYLE_STILL_ENABLED, ENCRYPTED_REASONING_DROPPED, NONE. Set file to the supplied patch filename. exact_values must contain, in order: tokens_in as a decimal integer string, tokens_cache_read as a string, then the correct total input counted once as a string. Obtain companion facts through tools; do not guess missing values.", + "source_ids": [ + "or-call-logs", + "promptfoo-python" + ], + "evidence_class": "synthetic mutation derived from cited upstream behavior", + "files": { + "usage.py": "def total_input(row):\n return row[\"tokens_in\"] + row[\"tokens_cache_read\"]\n" + }, + "companion_tool_input": { + "canary_packet_sha256": "8bf066389d8523cbfa1d3593d399a986083b6e875e652901f266469dce5a6d07", + "usage_row": { + "tokens_in": 1000, + "tokens_cache_read": 940 + } + }, + "expected": { + "defect_id": "CACHE_DOUBLE_COUNT", + "file": "usage.py", + "exact_values": [ + "1000", + "940", + "1000" + ] + }, + "oracle": "promptfoo native Python assertion: strict JSON parse, reject duplicate keys/nonfinite values, require exactly expected keys/types and byte-exact string values; exact object equality; malformed or missing output fails. Artifact-derived source facts, never model self-score.", + "schema": { + "type": "object", + "additionalProperties": false, + "required": [ + "defect_id", + "file", + "exact_values" + ], + "properties": { + "defect_id": { + "type": "string", + "enum": [ + "JSON_NUMBER_ROUNDING", + "MESSAGE_KEY_COLLISION", + "TOOL_TAIL_LOSS", + "CACHE_DOUBLE_COUNT", + "LEGACY_STYLE_STILL_ENABLED", + "ENCRYPTED_REASONING_DROPPED", + "NONE" + ] + }, + "file": { + "type": "string" + }, + "exact_values": { + "type": "array", + "items": { + "type": "string" + } + } + } + }, + "grader_controls": [ + { + "kind": "known-pass", + "output": "{\"defect_id\":\"CACHE_DOUBLE_COUNT\",\"file\":\"usage.py\",\"exact_values\":[\"1000\",\"940\",\"1000\"]}", + "expected_accept": true + }, + { + "kind": "known-fail", + "output": "{\"defect_id\":\"WRONG_CONTROL_VALUE\",\"file\":\"usage.py\",\"exact_values\":[\"1000\",\"940\",\"1000\"]}", + "expected_accept": false + }, + { + "kind": "malformed", + "output": "{\"unterminated\":", + "expected_accept": false + } + ] + }, + { + "id": "R-style", + "prompt": "Read the supplied patch file and companion tool evidence. Find the one planted defect under the frozen contract. Return only {\"defect_id\": string, \"file\": string, \"exact_values\": string[]}; do not edit files. Explain no additional speculative defects. Allowed defect_id vocabulary: JSON_NUMBER_ROUNDING, MESSAGE_KEY_COLLISION, TOOL_TAIL_LOSS, CACHE_DOUBLE_COUNT, LEGACY_STYLE_STILL_ENABLED, ENCRYPTED_REASONING_DROPPED, NONE. Set file to the supplied patch filename. exact_values must contain, in order: the resolved output style id, then its level. Obtain companion facts through tools; do not guess missing values.", + "source_ids": [ + "or-styles", + "promptfoo-python" + ], + "evidence_class": "synthetic mutation derived from cited upstream behavior", + "files": { + "styles.json": "{\"outputStyles\": [], \"cavemanOutputMode\": {\"enabled\": true, \"intensity\": \"lite\"}}\n" + }, + "companion_tool_input": { + "canary_packet_sha256": "8bf066389d8523cbfa1d3593d399a986083b6e875e652901f266469dce5a6d07", + "usage_row": { + "tokens_in": 1000, + "tokens_cache_read": 940 + } + }, + "expected": { + "defect_id": "LEGACY_STYLE_STILL_ENABLED", + "file": "styles.json", + "exact_values": [ + "terse-prose", + "lite" + ] + }, + "oracle": "promptfoo native Python assertion: strict JSON parse, reject duplicate keys/nonfinite values, require exactly expected keys/types and byte-exact string values; exact object equality; malformed or missing output fails. Artifact-derived source facts, never model self-score.", + "schema": { + "type": "object", + "additionalProperties": false, + "required": [ + "defect_id", + "file", + "exact_values" + ], + "properties": { + "defect_id": { + "type": "string", + "enum": [ + "JSON_NUMBER_ROUNDING", + "MESSAGE_KEY_COLLISION", + "TOOL_TAIL_LOSS", + "CACHE_DOUBLE_COUNT", + "LEGACY_STYLE_STILL_ENABLED", + "ENCRYPTED_REASONING_DROPPED", + "NONE" + ] + }, + "file": { + "type": "string" + }, + "exact_values": { + "type": "array", + "items": { + "type": "string" + } + } + } + }, + "grader_controls": [ + { + "kind": "known-pass", + "output": "{\"defect_id\":\"LEGACY_STYLE_STILL_ENABLED\",\"file\":\"styles.json\",\"exact_values\":[\"terse-prose\",\"lite\"]}", + "expected_accept": true + }, + { + "kind": "known-fail", + "output": "{\"defect_id\":\"WRONG_CONTROL_VALUE\",\"file\":\"styles.json\",\"exact_values\":[\"terse-prose\",\"lite\"]}", + "expected_accept": false + }, + { + "kind": "malformed", + "output": "{\"unterminated\":", + "expected_accept": false + } + ] + }, + { + "id": "R-replay", + "prompt": "Read the supplied patch file and companion tool evidence. Find the one planted defect under the frozen contract. Return only {\"defect_id\": string, \"file\": string, \"exact_values\": string[]}; do not edit files. Explain no additional speculative defects. Allowed defect_id vocabulary: JSON_NUMBER_ROUNDING, MESSAGE_KEY_COLLISION, TOOL_TAIL_LOSS, CACHE_DOUBLE_COUNT, LEGACY_STYLE_STILL_ENABLED, ENCRYPTED_REASONING_DROPPED, NONE. Set file to the supplied patch filename. exact_values must contain, in order: the dropped item type, then the field name containing its opaque encrypted bytes. Obtain companion facts through tools; do not guess missing values.", + "source_ids": [ + "or-codex-transport", + "promptfoo-python" + ], + "evidence_class": "synthetic mutation derived from cited upstream behavior", + "files": { + "replay.py": "def next_input(items):\n return [item for item in items if item[\"type\"] != \"reasoning\"]\n" + }, + "companion_tool_input": { + "canary_packet_sha256": "8bf066389d8523cbfa1d3593d399a986083b6e875e652901f266469dce5a6d07", + "usage_row": { + "tokens_in": 1000, + "tokens_cache_read": 940 + } + }, + "expected": { + "defect_id": "ENCRYPTED_REASONING_DROPPED", + "file": "replay.py", + "exact_values": [ + "reasoning", + "encrypted_content" + ] + }, + "oracle": "promptfoo native Python assertion: strict JSON parse, reject duplicate keys/nonfinite values, require exactly expected keys/types and byte-exact string values; exact object equality; malformed or missing output fails. Artifact-derived source facts, never model self-score.", + "schema": { + "type": "object", + "additionalProperties": false, + "required": [ + "defect_id", + "file", + "exact_values" + ], + "properties": { + "defect_id": { + "type": "string", + "enum": [ + "JSON_NUMBER_ROUNDING", + "MESSAGE_KEY_COLLISION", + "TOOL_TAIL_LOSS", + "CACHE_DOUBLE_COUNT", + "LEGACY_STYLE_STILL_ENABLED", + "ENCRYPTED_REASONING_DROPPED", + "NONE" + ] + }, + "file": { + "type": "string" + }, + "exact_values": { + "type": "array", + "items": { + "type": "string" + } + } + } + }, + "grader_controls": [ + { + "kind": "known-pass", + "output": "{\"defect_id\":\"ENCRYPTED_REASONING_DROPPED\",\"file\":\"replay.py\",\"exact_values\":[\"reasoning\",\"encrypted_content\"]}", + "expected_accept": true + }, + { + "kind": "known-fail", + "output": "{\"defect_id\":\"WRONG_CONTROL_VALUE\",\"file\":\"replay.py\",\"exact_values\":[\"reasoning\",\"encrypted_content\"]}", + "expected_accept": false + }, + { + "kind": "malformed", + "output": "{\"unterminated\":", + "expected_accept": false + } + ] + } + ], + "controls": [ + { + "kind": "known-pass", + "input": { + "operation": "feed each complete expected object through full frozen adapter" + }, + "expected_accept": true + }, + { + "kind": "known-fail", + "input": { + "operation": "feed each expected object with its first required value changed; numeric canary uses 1234567890123456800 or 1.5" + }, + "expected_accept": false + }, + { + "kind": "malformed", + "input": { + "output_bytes": "{\"unterminated\":", + "also_reject": [ + "missing output", + "duplicate keys", + "NaN", + "wrong type", + "unknown extra key" + ] + }, + "expected_accept": false + } + ], + "canary_packet_sha256": "f06487b5af0fdab5a12dbb35cae00b4b6b0e6ec938311a961c9f5ea332446b43", + "shared_files": { + "evidence.json": "{\n \"source_class\": \"peer_report\",\n \"reported_cache_fraction\": 0.94,\n \"observed_cache_fraction\": null,\n \"live_zone_observed\": null\n}", + "usage.json": "{\n \"tokens_in\": 1000,\n \"tokens_cache_read\": 940,\n \"tokens_out\": 100,\n \"tokens_reasoning\": 25\n}", + "styles.json": "{\n \"outputStyles\": [\n {\n \"id\": \"terse-prose\",\n \"level\": \"full\"\n }\n ],\n \"cavemanOutputMode\": {\n \"enabled\": true,\n \"intensity\": \"lite\"\n }\n}", + "precedence.txt": "Explicit nonempty outputStyles wins; otherwise enabled cavemanOutputMode maps to terse-prose at its intensity. Source: outputStyles/backCompat.ts:13-28 at a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "numeric.json": "{\n \"integer\": 1234567890123456711,\n \"decimal\": 1.50,\n \"negative_decimal\": -0.0100,\n \"pin\": \"dd6e9607e\",\n \"sha256\": \"039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d\"\n}", + "long.txt": "row 000: preserve source identity and its exact lexical value.\nrow 001: preserve source identity and its exact lexical value.\nrow 002: preserve source identity and its exact lexical value.\nrow 003: preserve source identity and its exact lexical value.\nrow 004: preserve source identity and its exact lexical value.\nrow 005: preserve source identity and its exact lexical value.\nrow 006: preserve source identity and its exact lexical value.\nrow 007: preserve source identity and its exact lexical value.\nrow 008: preserve source identity and its exact lexical value.\nrow 009: preserve source identity and its exact lexical value.\nrow 010: preserve source identity and its exact lexical value.\nrow 011: preserve source identity and its exact lexical value.\nrow 012: preserve source identity and its exact lexical value.\nrow 013: preserve source identity and its exact lexical value.\nrow 014: preserve source identity and its exact lexical value.\nrow 015: preserve source identity and its exact lexical value.\nrow 016: preserve source identity and its exact lexical value.\nrow 017: preserve source identity and its exact lexical value.\nrow 018: preserve source identity and its exact lexical value.\nrow 019: preserve source identity and its exact lexical value.\nrow 020: preserve source identity and its exact lexical value.\nrow 021: preserve source identity and its exact lexical value.\nrow 022: preserve source identity and its exact lexical value.\nrow 023: preserve source identity and its exact lexical value.\nrow 024: preserve source identity and its exact lexical value.\nrow 025: preserve source identity and its exact lexical value.\nrow 026: preserve source identity and its exact lexical value.\nrow 027: preserve source identity and its exact lexical value.\nrow 028: preserve source identity and its exact lexical value.\nrow 029: preserve source identity and its exact lexical value.\nrow 030: preserve source identity and its exact lexical value.\nrow 031: preserve source identity and its exact lexical value.\nrow 032: preserve source identity and its exact lexical value.\nrow 033: preserve source identity and its exact lexical value.\nrow 034: preserve source identity and its exact lexical value.\nrow 035: preserve source identity and its exact lexical value.\nrow 036: preserve source identity and its exact lexical value.\nrow 037: preserve source identity and its exact lexical value.\nrow 038: preserve source identity and its exact lexical value.\nrow 039: preserve source identity and its exact lexical value.\nrow 040: preserve source identity and its exact lexical value.\nrow 041: preserve source identity and its exact lexical value.\nrow 042: preserve source identity and its exact lexical value.\nrow 043: preserve source identity and its exact lexical value.\nrow 044: preserve source identity and its exact lexical value.\nrow 045: preserve source identity and its exact lexical value.\nrow 046: preserve source identity and its exact lexical value.\nrow 047: preserve source identity and its exact lexical value.\nrow 048: preserve source identity and its exact lexical value.\nrow 049: preserve source identity and its exact lexical value.\nrow 050: preserve source identity and its exact lexical value.\nrow 051: preserve source identity and its exact lexical value.\nrow 052: preserve source identity and its exact lexical value.\nrow 053: preserve source identity and its exact lexical value.\nrow 054: preserve source identity and its exact lexical value.\nrow 055: preserve source identity and its exact lexical value.\nrow 056: preserve source identity and its exact lexical value.\nrow 057: preserve source identity and its exact lexical value.\nrow 058: preserve source identity and its exact lexical value.\nrow 059: preserve source identity and its exact lexical value.\nrow 060: preserve source identity and its exact lexical value.\nrow 061: preserve source identity and its exact lexical value.\nrow 062: preserve source identity and its exact lexical value.\nrow 063: preserve source identity and its exact lexical value.\nTAIL_PIN=dd6e9607e\nTAIL_SHA256=039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d\nTAIL_DECIMAL=1.50\n", + "replay.json": "[{\"type\":\"reasoning\",\"encrypted_content\":\"SYNTHETIC_FIXTURE_NOT_REAL_REASONING\"}]\n", + "artifact.json": "{\n \"pin\": \"dd6e9607e\",\n \"sha256\": \"039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d\"\n}\n" + }, + "delivery": "Native Codex task + same-session three-user-turn tool canary; promptfoo grades retained outputs offline. Upstream task verifiers retain their own separate rewards.", + "scope": "Frozen finite task set only; no population-wide or arbitrary repository role claim." + }, + "sha256": "336a23603b7c31e0f5c8b6a2f6f49f9b27ab77f79c12d3ab7df5ab016fcc6198", + "materialized_runtime_sha256": null, + "native_controls_result": null + }, + "researcher": { + "packet": { + "role": "researcher", + "tasks": [ + { + "id": "S-number", + "prompt": "Read numeric.json through the shell tool. Transcribe number lexemes as JSON strings without parsing them as floats. Return only the required JSON object.", + "source_ids": [ + "or-minify", + "promptfoo-python" + ], + "input_refs": [ + "canaries.packet.numeric_tool_output", + "canaries.packet.long_tool_output", + "task_sets.researcher.packet.shared_files" + ], + "expected": { + "integer": "1234567890123456711", + "decimal": "1.50", + "negative_decimal": "-0.0100" + }, + "schema": { + "type": "object", + "additionalProperties": false, + "required": [ + "integer", + "decimal", + "negative_decimal" + ], + "properties": { + "integer": { + "type": "string" + }, + "decimal": { + "type": "string" + }, + "negative_decimal": { + "type": "string" + } + } + }, + "oracle": "promptfoo native schema assertion plus Python exact typed-object comparison; retry once with the identical fixed schema reminder, retain initial failure and both outputs. Success requires first response correct for the primary endpoint.", + "grader_controls": [ + { + "kind": "known-pass", + "output": "{\"integer\":\"1234567890123456711\",\"decimal\":\"1.50\",\"negative_decimal\":\"-0.0100\"}", + "expected_accept": true + }, + { + "kind": "known-fail", + "output": "{\"integer\":\"WRONG_CONTROL_VALUE\",\"decimal\":\"1.50\",\"negative_decimal\":\"-0.0100\"}", + "expected_accept": false + }, + { + "kind": "malformed", + "output": "{\"unterminated\":", + "expected_accept": false + } + ] + }, + { + "id": "S-pin", + "prompt": "Read the artifact identity from tool evidence and return its exact pin and sha256 strings. Return only the required JSON object.", + "source_ids": [ + "or-dedup", + "promptfoo-python" + ], + "input_refs": [ + "canaries.packet.numeric_tool_output", + "canaries.packet.long_tool_output", + "task_sets.researcher.packet.shared_files" + ], + "expected": { + "pin": "dd6e9607e", + "sha256": "039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d" + }, + "schema": { + "type": "object", + "additionalProperties": false, + "required": [ + "pin", + "sha256" + ], + "properties": { + "pin": { + "type": "string" + }, + "sha256": { + "type": "string" + } + } + }, + "oracle": "promptfoo native schema assertion plus Python exact typed-object comparison; retry once with the identical fixed schema reminder, retain initial failure and both outputs. Success requires first response correct for the primary endpoint.", + "grader_controls": [ + { + "kind": "known-pass", + "output": "{\"pin\":\"dd6e9607e\",\"sha256\":\"039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d\"}", + "expected_accept": true + }, + { + "kind": "known-fail", + "output": "{\"pin\":\"WRONG_CONTROL_VALUE\",\"sha256\":\"039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d\"}", + "expected_accept": false + }, + { + "kind": "malformed", + "output": "{\"unterminated\":", + "expected_accept": false + } + ] + }, + { + "id": "S-tail", + "prompt": "Read long.txt through a tool and return the three TAIL fields as lowercase JSON keys, retaining exact value bytes. Return only the required JSON object.", + "source_ids": [ + "or-lite", + "promptfoo-python" + ], + "input_refs": [ + "canaries.packet.numeric_tool_output", + "canaries.packet.long_tool_output", + "task_sets.researcher.packet.shared_files" + ], + "expected": { + "tail_pin": "dd6e9607e", + "tail_sha256": "039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d", + "tail_decimal": "1.50" + }, + "schema": { + "type": "object", + "additionalProperties": false, + "required": [ + "tail_pin", + "tail_sha256", + "tail_decimal" + ], + "properties": { + "tail_pin": { + "type": "string" + }, + "tail_sha256": { + "type": "string" + }, + "tail_decimal": { + "type": "string" + } + } + }, + "oracle": "promptfoo native schema assertion plus Python exact typed-object comparison; retry once with the identical fixed schema reminder, retain initial failure and both outputs. Success requires first response correct for the primary endpoint.", + "grader_controls": [ + { + "kind": "known-pass", + "output": "{\"tail_pin\":\"dd6e9607e\",\"tail_sha256\":\"039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d\",\"tail_decimal\":\"1.50\"}", + "expected_accept": true + }, + { + "kind": "known-fail", + "output": "{\"tail_pin\":\"WRONG_CONTROL_VALUE\",\"tail_sha256\":\"039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d\",\"tail_decimal\":\"1.50\"}", + "expected_accept": false + }, + { + "kind": "malformed", + "output": "{\"unterminated\":", + "expected_accept": false + } + ] + }, + { + "id": "S-null", + "prompt": "Read evidence.json. Return cache_fraction and live_zone_observed as null because no live observations exist; source_class must retain its literal value. Return only the required JSON object.", + "source_ids": [ + "or-live-zone", + "promptfoo-python" + ], + "input_refs": [ + "canaries.packet.numeric_tool_output", + "canaries.packet.long_tool_output", + "task_sets.researcher.packet.shared_files" + ], + "expected": { + "cache_fraction": null, + "live_zone_observed": null, + "source_class": "peer_report" + }, + "schema": { + "type": "object", + "additionalProperties": false, + "required": [ + "cache_fraction", + "live_zone_observed", + "source_class" + ], + "properties": { + "cache_fraction": { + "type": "null" + }, + "live_zone_observed": { + "type": "null" + }, + "source_class": { + "type": "string" + } + } + }, + "oracle": "promptfoo native schema assertion plus Python exact typed-object comparison; retry once with the identical fixed schema reminder, retain initial failure and both outputs. Success requires first response correct for the primary endpoint.", + "grader_controls": [ + { + "kind": "known-pass", + "output": "{\"cache_fraction\":null,\"live_zone_observed\":null,\"source_class\":\"peer_report\"}", + "expected_accept": true + }, + { + "kind": "known-fail", + "output": "{\"cache_fraction\":\"WRONG_CONTROL_VALUE\",\"live_zone_observed\":null,\"source_class\":\"peer_report\"}", + "expected_accept": false + }, + { + "kind": "malformed", + "output": "{\"unterminated\":", + "expected_accept": false + } + ] + }, + { + "id": "S-accounting", + "prompt": "Read usage.json. Return integer totals; cache reads are a subset of input and reasoning is a subset of output, each counted once. Return only the required JSON object.", + "source_ids": [ + "or-call-logs", + "promptfoo-python" + ], + "input_refs": [ + "canaries.packet.numeric_tool_output", + "canaries.packet.long_tool_output", + "task_sets.researcher.packet.shared_files" + ], + "expected": { + "input_total": 1000, + "cache_read_subset": 940, + "uncached_input": 60, + "reasoning_subset": 25, + "output_total": 100, + "total_tokens": 1100 + }, + "schema": { + "type": "object", + "additionalProperties": false, + "required": [ + "input_total", + "cache_read_subset", + "uncached_input", + "reasoning_subset", + "output_total", + "total_tokens" + ], + "properties": { + "input_total": { + "type": "integer" + }, + "cache_read_subset": { + "type": "integer" + }, + "uncached_input": { + "type": "integer" + }, + "reasoning_subset": { + "type": "integer" + }, + "output_total": { + "type": "integer" + }, + "total_tokens": { + "type": "integer" + } + } + }, + "oracle": "promptfoo native schema assertion plus Python exact typed-object comparison; retry once with the identical fixed schema reminder, retain initial failure and both outputs. Success requires first response correct for the primary endpoint.", + "grader_controls": [ + { + "kind": "known-pass", + "output": "{\"input_total\":1000,\"cache_read_subset\":940,\"uncached_input\":60,\"reasoning_subset\":25,\"output_total\":100,\"total_tokens\":1100}", + "expected_accept": true + }, + { + "kind": "known-fail", + "output": "{\"input_total\":\"WRONG_CONTROL_VALUE\",\"cache_read_subset\":940,\"uncached_input\":60,\"reasoning_subset\":25,\"output_total\":100,\"total_tokens\":1100}", + "expected_accept": false + }, + { + "kind": "malformed", + "output": "{\"unterminated\":", + "expected_accept": false + } + ] + }, + { + "id": "S-style", + "prompt": "Read styles.json and the pinned precedence excerpt. Report the selected style/level and whether legacy caveman adds another style. Return only the required JSON object.", + "source_ids": [ + "or-styles", + "promptfoo-python" + ], + "input_refs": [ + "canaries.packet.numeric_tool_output", + "canaries.packet.long_tool_output", + "task_sets.researcher.packet.shared_files" + ], + "expected": { + "selected_style": "terse-prose", + "selected_level": "full", + "legacy_added": false + }, + "schema": { + "type": "object", + "additionalProperties": false, + "required": [ + "selected_style", + "selected_level", + "legacy_added" + ], + "properties": { + "selected_style": { + "type": "string" + }, + "selected_level": { + "type": "string" + }, + "legacy_added": { + "type": "boolean" + } + } + }, + "oracle": "promptfoo native schema assertion plus Python exact typed-object comparison; retry once with the identical fixed schema reminder, retain initial failure and both outputs. Success requires first response correct for the primary endpoint.", + "grader_controls": [ + { + "kind": "known-pass", + "output": "{\"selected_style\":\"terse-prose\",\"selected_level\":\"full\",\"legacy_added\":false}", + "expected_accept": true + }, + { + "kind": "known-fail", + "output": "{\"selected_style\":\"WRONG_CONTROL_VALUE\",\"selected_level\":\"full\",\"legacy_added\":false}", + "expected_accept": false + }, + { + "kind": "malformed", + "output": "{\"unterminated\":", + "expected_accept": false + } + ] + } + ], + "controls": [ + { + "kind": "known-pass", + "input": { + "operation": "feed each complete expected object through full frozen adapter" + }, + "expected_accept": true + }, + { + "kind": "known-fail", + "input": { + "operation": "feed each expected object with its first required value changed; numeric canary uses 1234567890123456800 or 1.5" + }, + "expected_accept": false + }, + { + "kind": "malformed", + "input": { + "output_bytes": "{\"unterminated\":", + "also_reject": [ + "missing output", + "duplicate keys", + "NaN", + "wrong type", + "unknown extra key" + ] + }, + "expected_accept": false + } + ], + "canary_packet_sha256": "f06487b5af0fdab5a12dbb35cae00b4b6b0e6ec938311a961c9f5ea332446b43", + "shared_files": { + "evidence.json": "{\n \"source_class\": \"peer_report\",\n \"reported_cache_fraction\": 0.94,\n \"observed_cache_fraction\": null,\n \"live_zone_observed\": null\n}", + "usage.json": "{\n \"tokens_in\": 1000,\n \"tokens_cache_read\": 940,\n \"tokens_out\": 100,\n \"tokens_reasoning\": 25\n}", + "styles.json": "{\n \"outputStyles\": [\n {\n \"id\": \"terse-prose\",\n \"level\": \"full\"\n }\n ],\n \"cavemanOutputMode\": {\n \"enabled\": true,\n \"intensity\": \"lite\"\n }\n}", + "precedence.txt": "Explicit nonempty outputStyles wins; otherwise enabled cavemanOutputMode maps to terse-prose at its intensity. Source: outputStyles/backCompat.ts:13-28 at a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3", + "numeric.json": "{\n \"integer\": 1234567890123456711,\n \"decimal\": 1.50,\n \"negative_decimal\": -0.0100,\n \"pin\": \"dd6e9607e\",\n \"sha256\": \"039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d\"\n}", + "long.txt": "row 000: preserve source identity and its exact lexical value.\nrow 001: preserve source identity and its exact lexical value.\nrow 002: preserve source identity and its exact lexical value.\nrow 003: preserve source identity and its exact lexical value.\nrow 004: preserve source identity and its exact lexical value.\nrow 005: preserve source identity and its exact lexical value.\nrow 006: preserve source identity and its exact lexical value.\nrow 007: preserve source identity and its exact lexical value.\nrow 008: preserve source identity and its exact lexical value.\nrow 009: preserve source identity and its exact lexical value.\nrow 010: preserve source identity and its exact lexical value.\nrow 011: preserve source identity and its exact lexical value.\nrow 012: preserve source identity and its exact lexical value.\nrow 013: preserve source identity and its exact lexical value.\nrow 014: preserve source identity and its exact lexical value.\nrow 015: preserve source identity and its exact lexical value.\nrow 016: preserve source identity and its exact lexical value.\nrow 017: preserve source identity and its exact lexical value.\nrow 018: preserve source identity and its exact lexical value.\nrow 019: preserve source identity and its exact lexical value.\nrow 020: preserve source identity and its exact lexical value.\nrow 021: preserve source identity and its exact lexical value.\nrow 022: preserve source identity and its exact lexical value.\nrow 023: preserve source identity and its exact lexical value.\nrow 024: preserve source identity and its exact lexical value.\nrow 025: preserve source identity and its exact lexical value.\nrow 026: preserve source identity and its exact lexical value.\nrow 027: preserve source identity and its exact lexical value.\nrow 028: preserve source identity and its exact lexical value.\nrow 029: preserve source identity and its exact lexical value.\nrow 030: preserve source identity and its exact lexical value.\nrow 031: preserve source identity and its exact lexical value.\nrow 032: preserve source identity and its exact lexical value.\nrow 033: preserve source identity and its exact lexical value.\nrow 034: preserve source identity and its exact lexical value.\nrow 035: preserve source identity and its exact lexical value.\nrow 036: preserve source identity and its exact lexical value.\nrow 037: preserve source identity and its exact lexical value.\nrow 038: preserve source identity and its exact lexical value.\nrow 039: preserve source identity and its exact lexical value.\nrow 040: preserve source identity and its exact lexical value.\nrow 041: preserve source identity and its exact lexical value.\nrow 042: preserve source identity and its exact lexical value.\nrow 043: preserve source identity and its exact lexical value.\nrow 044: preserve source identity and its exact lexical value.\nrow 045: preserve source identity and its exact lexical value.\nrow 046: preserve source identity and its exact lexical value.\nrow 047: preserve source identity and its exact lexical value.\nrow 048: preserve source identity and its exact lexical value.\nrow 049: preserve source identity and its exact lexical value.\nrow 050: preserve source identity and its exact lexical value.\nrow 051: preserve source identity and its exact lexical value.\nrow 052: preserve source identity and its exact lexical value.\nrow 053: preserve source identity and its exact lexical value.\nrow 054: preserve source identity and its exact lexical value.\nrow 055: preserve source identity and its exact lexical value.\nrow 056: preserve source identity and its exact lexical value.\nrow 057: preserve source identity and its exact lexical value.\nrow 058: preserve source identity and its exact lexical value.\nrow 059: preserve source identity and its exact lexical value.\nrow 060: preserve source identity and its exact lexical value.\nrow 061: preserve source identity and its exact lexical value.\nrow 062: preserve source identity and its exact lexical value.\nrow 063: preserve source identity and its exact lexical value.\nTAIL_PIN=dd6e9607e\nTAIL_SHA256=039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d\nTAIL_DECIMAL=1.50\n", + "replay.json": "[{\"type\":\"reasoning\",\"encrypted_content\":\"SYNTHETIC_FIXTURE_NOT_REAL_REASONING\"}]\n", + "artifact.json": "{\n \"pin\": \"dd6e9607e\",\n \"sha256\": \"039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d\"\n}\n" + }, + "delivery": "Native Codex task + same-session three-user-turn tool canary; promptfoo grades retained outputs offline. Upstream task verifiers retain their own separate rewards.", + "scope": "Frozen finite task set only; no population-wide or arbitrary repository role claim." + }, + "sha256": "eda4967d6e0dfb1bf9d433c0fbf55c04c14a3603299d3937488d109daa07b056", + "materialized_runtime_sha256": null, + "native_controls_result": null + } + }, + "canaries": { + "packet": { + "roles": [ + "builder", + "reviewer", + "researcher" + ], + "evidence_class": "synthetic source-derived canaries", + "source_ids": [ + "or-minify", + "or-lite", + "or-dedup", + "or-styles" + ], + "numeric_tool_output": "{\n \"integer\": 1234567890123456711,\n \"decimal\": 1.50,\n \"negative_decimal\": -0.0100,\n \"pin\": \"dd6e9607e\",\n \"sha256\": \"039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d\"\n}", + "long_tool_output": "row 000: preserve source identity and its exact lexical value.\nrow 001: preserve source identity and its exact lexical value.\nrow 002: preserve source identity and its exact lexical value.\nrow 003: preserve source identity and its exact lexical value.\nrow 004: preserve source identity and its exact lexical value.\nrow 005: preserve source identity and its exact lexical value.\nrow 006: preserve source identity and its exact lexical value.\nrow 007: preserve source identity and its exact lexical value.\nrow 008: preserve source identity and its exact lexical value.\nrow 009: preserve source identity and its exact lexical value.\nrow 010: preserve source identity and its exact lexical value.\nrow 011: preserve source identity and its exact lexical value.\nrow 012: preserve source identity and its exact lexical value.\nrow 013: preserve source identity and its exact lexical value.\nrow 014: preserve source identity and its exact lexical value.\nrow 015: preserve source identity and its exact lexical value.\nrow 016: preserve source identity and its exact lexical value.\nrow 017: preserve source identity and its exact lexical value.\nrow 018: preserve source identity and its exact lexical value.\nrow 019: preserve source identity and its exact lexical value.\nrow 020: preserve source identity and its exact lexical value.\nrow 021: preserve source identity and its exact lexical value.\nrow 022: preserve source identity and its exact lexical value.\nrow 023: preserve source identity and its exact lexical value.\nrow 024: preserve source identity and its exact lexical value.\nrow 025: preserve source identity and its exact lexical value.\nrow 026: preserve source identity and its exact lexical value.\nrow 027: preserve source identity and its exact lexical value.\nrow 028: preserve source identity and its exact lexical value.\nrow 029: preserve source identity and its exact lexical value.\nrow 030: preserve source identity and its exact lexical value.\nrow 031: preserve source identity and its exact lexical value.\nrow 032: preserve source identity and its exact lexical value.\nrow 033: preserve source identity and its exact lexical value.\nrow 034: preserve source identity and its exact lexical value.\nrow 035: preserve source identity and its exact lexical value.\nrow 036: preserve source identity and its exact lexical value.\nrow 037: preserve source identity and its exact lexical value.\nrow 038: preserve source identity and its exact lexical value.\nrow 039: preserve source identity and its exact lexical value.\nrow 040: preserve source identity and its exact lexical value.\nrow 041: preserve source identity and its exact lexical value.\nrow 042: preserve source identity and its exact lexical value.\nrow 043: preserve source identity and its exact lexical value.\nrow 044: preserve source identity and its exact lexical value.\nrow 045: preserve source identity and its exact lexical value.\nrow 046: preserve source identity and its exact lexical value.\nrow 047: preserve source identity and its exact lexical value.\nrow 048: preserve source identity and its exact lexical value.\nrow 049: preserve source identity and its exact lexical value.\nrow 050: preserve source identity and its exact lexical value.\nrow 051: preserve source identity and its exact lexical value.\nrow 052: preserve source identity and its exact lexical value.\nrow 053: preserve source identity and its exact lexical value.\nrow 054: preserve source identity and its exact lexical value.\nrow 055: preserve source identity and its exact lexical value.\nrow 056: preserve source identity and its exact lexical value.\nrow 057: preserve source identity and its exact lexical value.\nrow 058: preserve source identity and its exact lexical value.\nrow 059: preserve source identity and its exact lexical value.\nrow 060: preserve source identity and its exact lexical value.\nrow 061: preserve source identity and its exact lexical value.\nrow 062: preserve source identity and its exact lexical value.\nrow 063: preserve source identity and its exact lexical value.\nTAIL_PIN=dd6e9607e\nTAIL_SHA256=039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d\nTAIL_DECIMAL=1.50\n", + "expected": { + "integer_lexeme": "1234567890123456711", + "decimal_lexeme": "1.50", + "negative_decimal_lexeme": "-0.0100", + "pin": "dd6e9607e", + "sha256": "039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d", + "tail_pin": "dd6e9607e", + "tail_decimal": "1.50" + }, + "numeric_padding": "Compare whole native output bytes with minBytes=512, maxCandidateBytes=512 KiB and effective configured limits; whole-string parse and actual numeric rewrite are separate predicates. Text can be exact but minifier-unreachable.", + "dedup_probe": "Direct chat-format stimulus: message[0] multipart has padding part[0] and pin part[1] (composite key 2). Message[1] string owns block B at key 1; message[2] repeats B at key 2. Dedup replacement key 2 can rewrite message[0].content[1]. Assert that earlier pin part is unchanged. Native Responses input_text parts may never enter this text-part path.", + "dedup_messages": [ + { + "role": "user", + "content": [ + { + "type": "text", + "text": "PADDING_ONLY=alpha\nNo repeated block here." + }, + { + "type": "text", + "text": "EARLIER_USER_PIN=0f76651d\nPreserve this independent earlier user value exactly." + } + ] + }, + { + "role": "tool", + "content": "line 00: this repeated multi-line block identifies artifact alpha exactly.\nline 01: this repeated multi-line block identifies artifact alpha exactly.\nline 02: this repeated multi-line block identifies artifact alpha exactly.\nline 03: this repeated multi-line block identifies artifact alpha exactly.\nline 04: this repeated multi-line block identifies artifact alpha exactly.\nline 05: this repeated multi-line block identifies artifact alpha exactly.\nline 06: this repeated multi-line block identifies artifact alpha exactly.\nline 07: this repeated multi-line block identifies artifact alpha exactly.\nline 08: this repeated multi-line block identifies artifact alpha exactly.\nline 09: this repeated multi-line block identifies artifact alpha exactly.\nline 10: this repeated multi-line block identifies artifact alpha exactly.\nline 11: this repeated multi-line block identifies artifact alpha exactly." + }, + { + "role": "tool", + "content": "line 00: this repeated multi-line block identifies artifact alpha exactly.\nline 01: this repeated multi-line block identifies artifact alpha exactly.\nline 02: this repeated multi-line block identifies artifact alpha exactly.\nline 03: this repeated multi-line block identifies artifact alpha exactly.\nline 04: this repeated multi-line block identifies artifact alpha exactly.\nline 05: this repeated multi-line block identifies artifact alpha exactly.\nline 06: this repeated multi-line block identifies artifact alpha exactly.\nline 07: this repeated multi-line block identifies artifact alpha exactly.\nline 08: this repeated multi-line block identifies artifact alpha exactly.\nline 09: this repeated multi-line block identifies artifact alpha exactly.\nline 10: this repeated multi-line block identifies artifact alpha exactly.\nline 11: this repeated multi-line block identifies artifact alpha exactly." + } + ], + "output_contract": "Return a UTF-8 JSON object of the expected string fields and complete requested artifact, with no numeric coercion, ellipsis or missing patch body. No expected answer is shown to the model.", + "qualification": "Run pinned native engine tests and local stimuli separately before qualification model cells. Record per shape: original bytes hash, adapter eligibility, whole-string parse, application, before/after equality and reason. Report unreachable on this lane where appropriate. No fabricated native execution or inferred pass from an unapplied engine.", + "numeric_engine_tool_output": "{\n \"integer\": 1234567890123456711,\n \"decimal\": 1.50,\n \"negative_decimal\": -0.0100,\n \"pin\": \"dd6e9607e\",\n \"sha256\": \"039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d\",\n \"padding\": \"row 000: preserve source identity and its exact lexical value.\\nrow 001: preserve source identity and its exact lexical value.\\nrow 002: preserve source identity and its exact lexical value.\\nrow 003: preserve source identity and its exact lexical value.\\nrow 004: preserve source identity and its exact lexical value.\\nrow 005: preserve source identity and its exact lexical value.\\nrow 006: preserve source identity and its exact lexical value.\\nrow 007: preserve source identity and its exact lexical value.\\nrow 008: preserve source identity and its exact lexical value.\\nrow 009: preserve source identity and its exact lexical value.\\nrow 010: preserve source identity and its exact lexical value.\\nrow 011: preserve source identity and its exact lexical value.\\nrow 012: preserve source identity and its exact lexical value.\\nrow 013: preserve source identity and its exact lexical value.\\nrow 014: preserve source identity and its exact lexical value.\\nrow 015: preserve source identity and its exact lexical value.\\nrow 016: preserve source identity and its exact lexical value.\\nrow 017: preserve source identity and its exact lexical value.\\nrow 018: preserve source identity and its exact lexical value.\\nrow 019: preserve source identity and its exact lexical value.\\nrow 020: preserve source identity and its exact lexical value.\\nrow 021: preserve source identity and its exact lexical value.\\nrow 022: preserve source identity and its exact lexical value.\\nrow 023: preserve source identity and its exact lexical value.\\nrow 024: preserve source identity and its exact lexical value.\\nrow 025: preserve source identity and its exact lexical value.\\nrow 026: preserve source identity and its exact lexical value.\\nrow 027: preserve source identity and its exact lexical value.\\nrow 028: preserve source identity and its exact lexical value.\\nrow 029: preserve source identity and its exact lexical value.\\nrow 030: preserve source identity and its exact lexical value.\\nrow 031: preserve source identity and its exact lexical value.\\nrow 032: preserve source identity and its exact lexical value.\\nrow 033: preserve source identity and its exact lexical value.\\nrow 034: preserve source identity and its exact lexical value.\\nrow 035: preserve source identity and its exact lexical value.\\nrow 036: preserve source identity and its exact lexical value.\\nrow 037: preserve source identity and its exact lexical value.\\nrow 038: preserve source identity and its exact lexical value.\\nrow 039: preserve source identity and its exact lexical value.\\nrow 040: preserve source identity and its exact lexical value.\\nrow 041: preserve source identity and its exact lexical value.\\nrow 042: preserve source identity and its exact lexical value.\\nrow 043: preserve source identity and its exact lexical value.\\nrow 044: preserve source identity and its exact lexical value.\\nrow 045: preserve source identity and its exact lexical value.\\nrow 046: preserve source identity and its exact lexical value.\\nrow 047: preserve source identity and its exact lexical value.\\nrow 048: preserve source identity and its exact lexical value.\\nrow 049: preserve source identity and its exact lexical value.\\nrow 050: preserve source identity and its exact lexical value.\\nrow 051: preserve source identity and its exact lexical value.\\nrow 052: preserve source identity and its exact lexical value.\\nrow 053: preserve source identity and its exact lexical value.\\nrow 054: preserve source identity and its exact lexical value.\\nrow 055: preserve source identity and its exact lexical value.\\nrow 056: preserve source identity and its exact lexical value.\\nrow 057: preserve source identity and its exact lexical value.\\nrow 058: preserve source identity and its exact lexical value.\\nrow 059: preserve source identity and its exact lexical value.\\nrow 060: preserve source identity and its exact lexical value.\\nrow 061: preserve source identity and its exact lexical value.\\nrow 062: preserve source identity and its exact lexical value.\\nrow 063: preserve source identity and its exact lexical value.\\nTAIL_PIN=dd6e9607e\\nTAIL_SHA256=039df6938ecc0adc0f5995f4bce10161a66318712c5f0856507c6dc17f6a675d\\nTAIL_DECIMAL=1.50\\n\"\n}", + "scheduled_tools": [ + "Turn 1: original task's first native step plus numeric/long evidence reads in the explicitly labelled local three-turn integration wrapper; retain original instruction and verifier bytes.", + "Turn 2: read fresh next-turn.txt (literal content: turn=2\\n); inspect prior tool output through native continuation.", + "Turn 3: read fresh final-turn.txt (literal content: turn=3\\n); answer using original numeric/tail/earlier-user evidence." + ], + "numeric_shapes": [ + { + "shape": "shell", + "stimulus_ref": "numeric_engine_tool_output", + "source_ids": [ + "codex-tool-shape", + "or-minify", + "or-adapter" + ], + "qualification": "Emit numeric JSON through native unified_exec; capture complete model-visible string including exit/time/chunk headers. Whole-string JSON parse is expected to fail by source. Keep exact-value scoring; label minifier unreachable on this shape, never strip headers to force eligibility.", + "native_result": null, + "eligible": null, + "whole_string_json_parse": null + }, + { + "shape": "mcp-text", + "stimulus_ref": "numeric_engine_tool_output", + "source_ids": [ + "codex-tool-shape", + "or-minify", + "or-adapter" + ], + "qualification": "MCP text content contains the numeric JSON bytes. Observe native function_call_output representation and protected-tool metadata; parse the entire eligible text, measure min/max byte thresholds and actual rewrite before claiming minifier coverage.", + "native_result": null, + "eligible": null, + "whole_string_json_parse": null + }, + { + "shape": "mcp-structuredContent", + "stimulus_ref": "numeric_engine_tool_output", + "source_ids": [ + "codex-tool-shape", + "or-minify", + "or-adapter" + ], + "qualification": "Supply lexical decimal/integer values as strings in structuredContent plus the raw numeric JSON as a text block. Numeric JSON serialization already loses scale before compression; separate native serialization loss from engine loss. Observe the model-visible representation and whole-string eligibility, do not assume structuredContent is delivered raw.", + "native_result": null, + "eligible": null, + "whole_string_json_parse": null + }, + { + "shape": "custom-tool", + "stimulus_ref": "numeric_engine_tool_output", + "source_ids": [ + "codex-tool-shape", + "or-minify", + "or-adapter" + ], + "qualification": "Use a source-supported native custom tool output carrying the same raw text; retain real matching call_id and custom_tool_call_output. If this runner cannot expose the tool or adapter eligibility rejects it, report unreachable on this lane and keep the engine coverage result open.", + "native_result": null, + "eligible": null, + "whole_string_json_parse": null + } + ], + "dedup_assert_unchanged": "messages[0].content[1].text", + "dedup_native_reachability": { + "result": null, + "source_ids": [ + "or-dedup", + "or-adapter", + "or-strategy" + ], + "requirement": "Record source and target wire formats, actual multipart part type, engine stage and applied telemetry before interpreting direct Chat reproduction as native-lane coverage. Source suggests input_text does not meet type=text condition; live reachability is unproven." + }, + "upstream_core_boundary": "Additional canary instructions/steps are local integration additions. Original native example commands/verifiers remain unchanged but the combined trial is not an unchanged upstream example acceptance run." + }, + "sha256": "8bf066389d8523cbfa1d3593d399a986083b6e875e652901f266469dce5a6d07", + "live_result": null + }, + "oracle_contract": { + "primary_success": "Recovered final core task success AND all exact record/canary/tool-roundtrip checks in every session including C. Engine not applied is neither a success override nor an outcome exclusion; application/reachability is a separate coverage gate.", + "grading": "Harbor native verifier for builder. promptfoo native schema/Python assertions on retained native outputs for reviewer/researcher. Locally materialized task-specific extraction/assertion adapters remain local integration glue, not unchanged upstream tests.", + "exact_comparison": "Reject invalid UTF-8, duplicate JSON keys, nonfinite numbers, missing/extra keys, type mismatches and any changed lexical record field. Compare string bytes, including decimal scale and pin/hash case. Independent original source bytes are oracle; no model self-reported pass.", + "required_negative_qualification": "Every adapter/verifier must accept its known-pass and reject known-fail, malformed and missing evidence through its full extraction path before any model cell. Exercise numeric rounded integer, 1.5, wrong pin/hash and tail removed. Empty execution/zero tests/exit0 without real reward fails.", + "adapters_sha256": null, + "native_control_results": null, + "blindness": "Sanitized cell labels hidden from grader and human adjudicator; deterministic canonical output comparisons. Unexpected but potentially correct outputs retained as failures under frozen rule, no post-hoc oracle edits.", + "score_exactness_when_engine_not_applied": true, + "engine_coverage_separate": true, + "recovery": "One fixed schema-reminder retry allowed within the same session/turn budget; retain initial failures and all event costs as secondary. Only failure still unresolved at the end enters operational_failure. No replacement session or retrospective oracle edit." + }, + "multi_turn": { + "user_turns": 3, + "tools_required_each_turn": true, + "aggregation_unit": "task/session/user_turn", + "same_native_session": true, + "sequence": [ + "Turn 1: original core task (create-file only for the native two-step task) plus canary tools in labelled integration wrapper. Capture full native output shape and original IDs.", + "Turn 2: native resume; append-content for the two-step task, otherwise recall prior evidence with a fresh tool action. Preserve real function/custom call IDs and any genuine encrypted reasoning.", + "Turn 3: require final exact record using earlier user provenance plus tail/numeric values, with another tool action; unchanged-file re-reads beyond scheduled tools count as recall." + ], + "required_checks": [ + "prompt_cache_key", + "session_headers", + "affinity_headers", + "store", + "include", + "encrypted_reasoning", + "turn_2_tool_roundtrip" + ], + "headers_to_trace": [ + "x-codex-session-id", + "x-session-id", + "x-omniroute-session", + "x-omniroute-session-id", + "session_id", + "x-codex-turn-state", + "x-omniroute-connection", + "X-Correlation-Id" + ], + "header_scope": "Separate session-affinity carriers, live-zone header/body fallback, connection pin and per-call correlation. Equal strings across distinct connection namespaces are not required. Never infer account spread from a pin header name.", + "wire_fields": { + "store": false, + "include": [ + "reasoning.encrypted_content" + ], + "prompt_cache_key": "stable unique session identity; separate across arms and repetitions" + }, + "field_acceptance": "Observe both ingress/egress hops privately. Assert stable cache/session identity within session, chosen account affinity unchanged, include remains effective, store=false accepted, encrypted reasoning bytes identical when replayed. If provider returns no encrypted item, mark that subcheck null and keep its gate open. No fabricated blob.", + "native_capability_gate": "Materialize native [[steps]] task with agent.resume_trajectory=true. Preserve same native thread across all three turns. Two-step builder uses its two original steps as turns 1/2 and adds exactly one record step; do not wrap three more turns around it. Hash wrappers; qualify each original step verifier plus final exact record.", + "native_route": "Harbor [[steps]]", + "resume_trajectory": true, + "task_turn_mapping": { + "B-hello-multi-step-simple": [ + "create-file", + "append-content", + "exact-record" + ], + "all_other_tasks": [ + "original-core-and-evidence", + "recall-with-fresh-tool", + "exact-record" + ] + }, + "affinity": { + "session_headers": [ + "x-codex-session-id", + "x-session-id", + "x-omniroute-session" + ], + "connection_pin": "x-omniroute-connection", + "account_spread_by_arm": true, + "body_derived_live_zone_fallback": true, + "source_ids": [ + "or-affinity", + "or-session-fallback", + "or-chat-core" + ], + "fallback_order": "Affinity: headers, metadata.session_id/sessionId, conversation_id, session_id, prompt_cache_key, first-input SHA256. Live zone: x-omniroute-session-id else generateSessionId(body, provider, connection). These are different mechanisms.", + "qualification": "Trace prompt_cache_key at both hops; retain equality booleans and salted account IDs per arm/turn. Compare account counts/concentration and cache-read ratios by arm and turn using pilot intervals. A dropped key or one-account collapse blocks confirmation until explained and frozen; occurrence remains unobserved." + } + }, + "hazards": [ + { + "id": "H1", + "source_ids": [ + "or-live-zone", + "or-chat-core" + ], + "verification_status": "Principal/session requirement verified in installed source; 94% is peer-reported historical context only; cache loss is a hypothesis.", + "checks": [ + "Keyed D/A cells with TTL 60 are confirmatory; verify live-zone principal/session and actual reuse.", + "Measure entry-gateway usage by turn with correlated rows, and per-arm account spread/cache reads.", + "Cache warmth follows actual within-session observations; no priming or claimed cold starts." + ], + "live_result": null + }, + { + "id": "H2", + "source_ids": [ + "or-minify", + "or-dedup", + "or-lite", + "or-adapter", + "or-strategy", + "or-engine-labels", + "or-caveman", + "or-relevance", + "or-aggressive", + "or-llmlingua", + "or-ultra" + ], + "verification_status": "Source-confirmed numeric normalization, key collision and truncation paths. Responses-to-lite path now traced; actual live settings/application remain unobserved.", + "checks": [ + "Compare number lexemes, pin/hash strings and >2000-character tail sentinels at wire, native transcript and deliverable boundaries.", + "Use role tool function_call_output for shell command plus protected read and unmatched call IDs; do not infer lite protection from codex-responses guards.", + "Direct Chat key 2 stimulus plus separate native Responses reachability: earlier message[0].content[1] must be unchanged; no-op outcomes still score exactness.", + "Record applied/no-eligible/skipped/error per engine; a no-op is not engine acceptance." + ], + "live_result": null + }, + { + "id": "H3", + "source_ids": [ + "or-chat-core", + "or-styles" + ], + "verification_status": "Configured styles run without allow-lossy when enabled/header not off. Explicit styles override legacy caveman; they are not additive.", + "checks": [ + "Cross default/all input plans with styles on/off. Off means outputStyles=[] AND cavemanOutputMode.enabled=false.", + "Grade all exact JSON record values and builder deliverables. Reviewer/researcher prose-style harmlessness is outside this task pack's scope.", + "Read back effective styles and applied telemetry each request; no cross-cell shared settings races." + ], + "live_result": null + }, + { + "id": "H4", + "source_ids": [ + "or-chat-core" + ], + "verification_status": "Header-off/reactive separation verified; safety passes gated by !nativeCodexPassthrough, so actual framework-route reachability is a live gate.", + "checks": [ + "Only C on 20128 is the clean control.", + "Exploratory bounded threshold probes below/above resolved proactive and final limits compare pre/post message inventories and hashes, including header off.", + "Reject any control with compression enabled or codex/* exclusion missing; retain safety-pass telemetry for framework cells." + ], + "live_result": null + }, + { + "id": "H5", + "source_ids": [ + "or-codex-transport", + "harbor-codex", + "promptfoo-responses", + "or-affinity-header", + "or-affinity", + "or-session-fallback", + "or-effort", + "or-correlation" + ], + "verification_status": "Some fields supported in source; end-to-end forwarding across sharedgw is not proven.", + "checks": [ + "Trace prompt_cache_key, session/affinity headers, store/include semantics and encrypted reasoning hashes at client->20129->20128->provider boundaries.", + "Turn 2 must receive real prior function_call_output matched to call_id and produce a valid next tool call and final answer.", + "Keep credentials/account-bound reasoning private; log only equality booleans and salted identifiers." + ], + "live_result": null + }, + { + "id": "H6", + "source_ids": [ + "role-policy" + ], + "verification_status": "Normative user role restriction, not an installed-package behavior.", + "checks": [ + "Builder-only early canary after deterministic and transport gates; reviewer/researcher trial outputs remain quarantined.", + "Review, judgment, verification and evidence routes stay 20128 until their own frozen quality and exact-value gates pass.", + "Record-output preservation is mandatory even for builders; task success alone cannot promote record-producing roles." + ], + "live_result": null + } + ], + "metrics": [ + { + "id": "task_success", + "definition": "Final recovered core success AND exact canary/record/tool-roundtrip success; initial failure and core score retained separately. C and no-op engine sessions still score exactness." + }, + { + "id": "apply_patch_failures", + "definition": "Count rejected/error apply_patch tool invocations per session; session-any-failure indicator. Inapplicable tasks remain null for this component; any malformed required patch fails task." + }, + { + "id": "verification_failures", + "definition": "Native nonzero verification invocations and missing/skipped/nonfinite rewards per session; intended starting red tests labelled separately; fixed protocol-required negative controls excluded from outcome denominator and charged to qualification." + }, + { + "id": "schema_retries", + "definition": "Count initial schema failures and up to one fixed-reminder retry as secondary; unresolved final schema failure enters both final success and unrecovered operational failure. All costs retained." + }, + { + "id": "tool_output_recall", + "definition": "An unrequested second read of the same normalized path, overlapping byte range and unchanged file SHA256 after a successful read; changed-file verification reads and mandated canary reads excluded. Count events and sessions, cite original/re-read call IDs; unresolved identity remains null." + }, + { + "id": "tokens_in", + "definition": "Final entry-gateway tokens_in once per correlation_id/row: C 20128, engines-on 20129." + }, + { + "id": "tokens_cache_read", + "definition": "Subset of tokens_in; report per-turn sum and fraction sum(cache)/sum(input), never mean of call fractions." + }, + { + "id": "tokens_reasoning", + "definition": "Entry-gateway output subset, never add to tokens_out." + }, + { + "id": "compression_savings", + "definition": "20129 per-call tokens_compressed joined by correlation_id; includes reactive compaction, diagnostic only. Isolated analytics optional, never additive." + }, + { + "id": "latency", + "definition": "Monotonic start-to-terminal wall seconds per task and user turn, includes tool, compression, verification and retry time; gateway duration diagnostic separate." + }, + { + "id": "cancellations", + "definition": "All status 499 rows attributable to owned attempts plus local aborted-without-row attempts; aborted row counts and aborted session counts separate. No automatic replacement; unknown usage stays null." + } + ], + "accounting": { + "usage_authority": { + "C": "20128.call_logs", + "D0": "20129.call_logs", + "D1": "20129.call_logs", + "A0": "20129.call_logs", + "A1": "20129.call_logs" + }, + "savings_authority": "20129.call_logs.tokens_compressed", + "add_savings_to_usage": false, + "cache_read_is_input_subset": true, + "reasoning_is_output_subset": true, + "required_columns": [ + "timestamp", + "path", + "status", + "model", + "reasoning_effort_requested", + "reasoning_effort_upstream", + "tokens_in", + "tokens_cache_read", + "tokens_reasoning", + "correlation_id", + "tokens_compressed" + ], + "required_additional_columns": [ + "id", + "tokens_out", + "duration" + ], + "output_column_verified_on_host": null, + "uncached_input_formula": "sum(tokens_in - tokens_cache_read)", + "total_billed_tokens_formula": "sum(tokens_in + tokens_out)", + "total_billed_tokens_scope": "Total provider-billed token positions, including cached input once, unweighted. This is not dollars; cached input can have a different price. Output is required because styles alter deliverables. Listed input/reasoning fields alone cannot determine the total.", + "billed_input_equivalent_formula": "sum(tokens_in - tokens_cache_read + cache_price_ratio * tokens_cache_read)", + "cache_price_ratio": null, + "economic_total_formula": "sum(uncached_input + cache_price_ratio * tokens_cache_read + output_price_ratio * tokens_out)", + "output_price_ratio": null, + "economic_rule": "Economic acceptance means robustness over the preregistered sensitivity rectangle, not inferred dollar savings. Unknown actual price ratios are not an automatic failure if every sensitivity corner passes. Any actual subscription meter requires its own qualification before making limit-consumption claims.", + "include_failed_retried_cancelled_attempts": true, + "join": "Join the response X-Correlation-Id captured on the client path to correlation_id at that arm's entry gateway; finalize/deduplicate by (gateway,row id). Keep private task/session/turn/request ordinals. All retries/tool continuations/failures count once. Exclude the engines-on chained 20128 usage row entirely. Never identify an experiment by model+time window or sum polling snapshots/hops.", + "negative_or_missing": "Require cache_read subset of input and reasoning subset of output. Malformed/negative usage is a telemetry failure. Null rows retain bounded intervals when justified by missing_usage; no zero-fill or automatic ineligibility for a bounded 499.", + "compression_snapshot": "Per-entry-row tokens_compressed is the latest recorded estimate and may describe reactive/context-fit compression instead of the selected pipeline: chatCore.ts L2164 overwrites the earlier L1916 estimate. It is not an additive total across engine stages, observed billing reduction or an engine-specific causal effect. Optional exclusive analytics deltas corroborate only; never add them to per-call savings or usage.", + "receipt_rows": [ + "role", + "task_id", + "cell", + "repetition", + "session_ordinal", + "user_turn", + "request_ordinal", + "start/end", + "status", + "model", + "reasoning_effort_requested", + "reasoning_effort_upstream", + "tokens_in", + "tokens_out", + "tokens_cache_read", + "tokens_reasoning", + "duration", + "native_artifact_sha256", + "verifier_result", + "initial_schema_result", + "exact_value_result", + "entry_gateway", + "correlation_id (salted public mapping)", + "request_body_reasoning_effort", + "tokens_compressed", + "engine_application", + "native_shape_eligibility", + "missing_usage_lower/upper" + ], + "zero_denominators": "No synthetic zero total, zero paired-session denominator or zero native tests. Ratios with missing/zero control usage are null and cannot qualify cost.", + "add_downstream_hop": false, + "correlation_capture": { + "header": "X-Correlation-Id", + "qualified": null, + "adapter_sha256": null, + "source_ids": [ + "or-correlation", + "codex-response-headers", + "mitmproxy-headers", + "mitmproxy-stream", + "mitmproxy-reverse" + ], + "mechanism": "Client-side mitmproxy 12.2.3 reverse listener on an owned loopback port, forwarding unchanged to the selected entry gateway. A minimal responseheaders adapter following upstream hook/stream examples exports only X-Correlation-Id, status, flow/request ordinal and assigned task/session/turn ordinal. It does not inspect, print or persist Authorization, other request headers, key values, bodies or full flow dumps. Supported client model/provider semantics remain native; the observer is local integration glue requiring materialization/hash/acceptance.", + "invocation": "mitmdump --mode reverse:http://127.0.0.1: --listen-host 127.0.0.1 --listen-port --set stream_large_bodies=1 -q -s ", + "client_capture": "Codex provider base_url points to that listener for all arms. Bind exactly one trial/turn at a time; every HTTP exchange has its own observer flow ID and returned correlation_id. Native --json output is not assumed to contain arbitrary response headers.", + "second_hop": "Effort-only downstream attribution requires a separate observed join from 20129 outgoing exchange to 20128 X-Correlation-Id. Qualify an identical header-only observer on sharedgw's outbound path, with explicit owned-session/flow identity and no unrelated lane traffic. No matching by timestamp/model. Until this second-hop join is demonstrated, effort acceptance remains open; chained usage is still excluded.", + "equivalence_gate": "Prove byte-identical model/body forwarding and unchanged SSE/tool/reasoning/cache behavior in all arms, safe redaction, error/499 capture and exact row join. Observer bytes/body hashes may be checked privately without copying keys. Install/pin/hash through upstream supported commands before qualification. No observer runs or host changes in repair." + }, + "effort": { + "field": "reasoning.effort", + "hops": [ + 20129, + 20128 + ], + "request_detail_route": "GET /api/usage/call-logs/", + "null_column_is_drift": false, + "source_ids": [ + "or-effort", + "or-detail-api" + ], + "rule": "At each joined hop inspect stored clientRequest/providerRequest body's reasoning.effort; print that field only. Hash/equality receipts stay value-free except literal effort=max. Freeze actual JSON selectors against returned detail schema. Columns reasoning_effort_requested/upstream corroborate only; a null column never means missing effort or drift. A missing stored body is unknown effort, an observed non-max body is drift." + }, + "sensitivity": { + "cache_ratio": [ + 0.0, + 1.0 + ], + "output_ratio": [ + 1.0, + 10.0 + ], + "actual_price_claim": false, + "source_ids": [ + "or-call-logs", + "scipy" + ], + "interpretation": "Protocol assumptions for subscription/limit-consumption sensitivity, not published GPT-6 prices. Evaluate all four corners of this rectangle using paired upper bounds and worst-case missing-usage intervals. Linear weighted-cost inequality reaches its extrema at corners; compare weighted candidate - 1.01*control as well as descriptive ratios. Report crossover weights; if ranking changes, no robust promotion or dollar claim. Actual price ratios and limit meter remain null.", + "stratification": "Whole session and each user-turn stratum 1/2/3; measured cache_read/input, no asserted cold/warm assignment." + }, + "missing_usage": { + "zero_fill": false, + "request_size_is_not_total_bound": true, + "worst_case_ranking_required": true, + "source_ids": [ + "or-call-logs", + "or-effort" + ], + "rule": "Keep null 499 usage null. For each missing row retain only nonsecret stored request byte/token size and the qualified provider serialization/context and output caps; derive [0, input_upper + output_upper] for total usage, including every attributable retry. Request size alone cannot bound output, reasoning or provider serialization overhead. Pilot counts cannot establish a hard upper bound. If these upper bounds are not supported, upper=null and cost remains inconclusive.", + "ranking": "Promote only if the candidate remains cheaper under candidate upper/control lower totals and all sensitivity corners, with uncertainty bounds. Known correctness failures remain outcomes, never replaced. Unbounded rows do not fabricate drift or trigger a routine-event stop; pause further spending when reserve cannot be bounded." + } + }, + "analysis": { + "tasks_per_role": 6, + "sessions_per_role_cell": null, + "total_confirmatory_sessions": null, + "success_ni_margin": 0.05, + "failure_ni_margin": 0.05, + "alpha": 0.05, + "upstream_functions": [ + "bootstrap", + "permutation_test", + "binomtest", + "multipletests" + ], + "estimand": "Per-role uniform six-task mixture. Same independent task draw paired across all five arms. Net success loss and excess unrecovered operational failure are session endpoints; repeated tasks do not broaden domain coverage.", + "failure_endpoint": "Unrecovered operational failure at final session termination after the fixed recovery allowance. Rejected patches, failed intermediate verification, schema retry and recovered request cancellations are secondary event counts; exact-record corruption still vetoes the cell.", + "noninferiority": "One-sided upper interval for paired net harm must be below +0.05 for both endpoints, using paired_net_test, then intersection-union candidate p and Holm across 12 role-candidate hypotheses. Harmful discordance alone is no longer the primary test.", + "assumptions": "Independent paired draws, stationarity over the sealed schedule and preserved within-draw dependencies. Analyze the finite task domain only. Pilot estimates and approximate bootstrap tests may fail calibration; then do not start confirmation. Material dependence, drift or broken randomization makes inference inconclusive.", + "multiplicity": { + "method": "holm", + "comparisons": 12, + "unit": "role-candidate intersection-union", + "family": "3 roles x 4 candidates. Candidate p=max(p_success,p_unrecovered_failure); both NI endpoints must pass. Missing hypotheses p=1; statsmodels multipletests(method=holm, alpha=.05). Intersection-union does not double the family merely for requiring both endpoints.", + "source_ids": [ + "intersection-union", + "statsmodels" + ] + }, + "bootstrap": "paired=True percentile 99999 resamples, PCG64 seed 20260927; shared complete draw indices across all cells/endpoints. NI uses paired_net_test. Cost/turn intervals attach all tools/retries to original session; simultaneous sensitivity guards use Bonferroni upper intervals over 12 candidates x 4 corners x 4 strata x 2 input/total metrics.", + "permutation": "scipy.stats.permutation_test permutation_type=samples, n_resamples=99999, alternative=less, rng seed=20260928, statistic mean(candidate_tokens-control_tokens) on paired complete session draws. Explanatory cost comparison only; nonsignificant equality is not non-inferiority.", + "degenerate_interval_guard": "Use the paired harmful-minus-beneficial exact fallback above when discordance count <20 or bootstrap is degenerate; unsupported ratios/nonfinite results remain inconclusive.", + "precision": "n and resource allocations remain null until the excluded qualification pilot and power/calibration simulation. No confirmatory peek, optional stopping, extra draws or alternative analysis after confirmation starts.", + "ordering": "For draw ordinal r use r mod 10 through [C,D0,A1,D1,A0], [D0,D1,C,A0,A1], [D1,A0,D0,A1,C], [A0,A1,D1,C,D0], [A1,C,A0,D0,D1] then their respective reversals. Add sorted role ordinal as cyclic schedule phase. Freeze realized confirmation schedule after pilot and before confirmatory data.", + "cache_design": "No priming sessions and no assigned cold/warm conditions. Fresh native identity per arm/draw, stable identity within all three turns; stratify cache and economic analysis by actual turn 1, 2, 3 plus session totals. Record actual account and cache-read fractions; no assumption of zero cache on turn 1 or warmth on turn 2.", + "planned_repetitions_per_role_cell": null, + "expected_repetitions_per_task_cell": null, + "sampling": "After excluded pilot and before confirmation, NumPy 2.4.0 Generator(PCG64(20260927 + sorted_role_ordinal)) draws n independent uniform task indices 0..5 with replacement. Freeze realized schedule and task counts/hash. All five cells use the same task draw; no cold/warm labels, balancing, replacement or post-outcome rerandomization.", + "scored_vs_priming": "Confirmatory sessions = 5 * sum(n_role), all scored; zero priming. Pilot/control/exploratory sessions charged separately.", + "actual_total_sessions_including_priming": null, + "operational_failure": { + "unrecovered_only": true, + "recovery": "One fixed schema-reminder retry within the same three-turn session. Native transport retries recorded at frozen limits; no trial replacement.", + "secondary": [ + "rejected patches", + "intermediate verification failures", + "schema retries", + "recovered cancellations", + "recall events" + ] + }, + "paired_net_test": { + "method": "scipy.stats.bootstrap", + "paired": true, + "statistic": "mean(control_success - candidate_success)", + "failure_statistic": "mean(candidate_unrecovered_failure - control_unrecovered_failure)", + "method_option": "percentile", + "n_resamples": 99999, + "seed": 20260927, + "p_rule": "For each endpoint D_i in {-1,0,1}, invert the one-sided percentile bound at margin .05: p=(1 + count(mean(D*) >= .05))/(99999+1), capped at 1. This is an approximate bootstrap NI test, not an exact finite-sample guarantee; calibrate before confirmation. At Holm rank r require corresponding upper (1-alpha/(13-r)) bound < .05. Preserve whole paired draws and shared C across candidates.", + "exact_fallback": "For fewer than 20 nonzero D_i or a zero-width/nonfinite interval, let h=count(D=+1) harmful and b=count(D=-1) beneficial. At candidate alpha a use scipy.stats.binomtest(h,n,alternative=less).proportion_ci(confidence_level=1-a/2,method=exact).high minus binomtest(b,n,alternative=greater).proportion_ci(confidence_level=1-a/2,method=exact).low. This Bonferroni upper bound on p_harmful-p_beneficial is conservative without independence of h and b. Require U < .05; invert monotonically over a in (0,1) for p, never substitute a harmful-only test. All-concordant data get a nonzero exact bound. If unsupported/nonfinite, p=1.", + "source_ids": [ + "scipy", + "scipy-exact", + "tango" + ], + "alternative_considered": "Tango 1998 score interval targets paired net difference. Not implemented here; pinned SciPy is chosen to avoid inventing a score solver. A failed pre-data bootstrap calibration leaves the gate open, not a post-data method switch.", + "alternative": "less" + }, + "priming_sessions": 0, + "cache_strata": [ + 1, + 2, + 3 + ], + "power_simulation": { + "target_power": 0.8, + "joint_paired_outcomes": true, + "margin_null_calibration": true, + "selected_n": null, + "results": null, + "n_grid": [ + 120, + 240, + 480, + 960, + 1920, + 3840 + ], + "replicates": 10000, + "seed": 20260929, + "source_ids": [ + "scipy", + "scipy-exact", + "statsmodels", + "intersection-union", + "numpy-pilot-simulation" + ], + "generation": "Pilot law: draw task uniformly then resample the full observed five-arm success/failure vector within that task (two pilot draws/task). For the harmless-candidate power law, uniformly permute all five arm labels in each sampled vector, making all margins equal while preserving success/failure coupling and multivariate dependence. Also retain the unsymmetrized empirical alternative. Sensitivity laws use success p in {.90,.95,.98,.99}, unrecovered failure f in {.01,.05,.10} only when f<=1-p, and correlation rho in {0,.5,.9}. Per draw, with probability rho use a shared uniform U for all five arms, otherwise independent U values; classify each arm as success if U=1-f, and recovered core failure otherwise. Infeasible p,f pairs are reported as mathematically impossible, never silently omitted. Generator(PCG64(seed)) and its random/integers/multinomial functions are the simulation primitives.", + "identical_candidate": "The exchangeable pilot resampling law and equal-threshold sensitivity laws have zero paired net harm. All endpoint and candidate labels remain coupled within each simulated draw; reuse shared C rather than drawing a fresh control for each contrast. Calculate the declared paired-net bootstrap/fallback and complete Holm family. Counts over the three-valued D distribution may use multinomial resampling only after equality to the pinned SciPy paired bootstrap is validated on boundary fixtures; no unqualified normal approximation.", + "calibration": "Repeat feasible sensitivity laws and empirical-pilot laws with each nonempty subset of candidates at one endpoint margin null: decrease success threshold p by .05, or increase failure f by .05 when feasible, holding the other endpoint within margin. Include all-null, single-null and mixed-null families, and common-control dependence within each role. Record all generated probabilities, infeasible cases and coverage gaps; if a required null law cannot be constructed from the pilot, qualification remains open. Require upper 95% Clopper-Pearson bound on family false promotion <=.055 (.005 Monte Carlo tolerance around .05). The coordinator must freeze and hash the executable law enumeration before simulation; no statistical result is asserted here.", + "n_rule": "Smallest n per role in the grid whose lower 95% exact binomial bound for jointly qualifying all four harmless candidates is >=.80 in the frozen pilot-based and sensitivity laws, and which passes margin-null calibration. Freeze n and all laws/output hashes before confirmation. If none qualifies or total resources are unavailable, remain unqualified; do not downsize n to fit budget." + } + }, + "decision_rule": { + "per_role": true, + "objective": "minimum total billed tokens among non-inferior cells", + "eligibility": "Complete matched cohort; qualified runner/profile/network/header joins, stored request-body effort=max at applicable hops, unchanged effective C, passed controls, zero exact record corruption, both NI endpoints via 12-hypothesis Holm, engine coverage labelled independently. Null effort columns do not disqualify; bounded missing usage uses worst-case cost ranking.", + "cost_rule": "All-attempt entry-gateway input+output totals on the same complete role schedule, including failed/retried/cancelled calls once, zero primes. Rank only after NI/exactness. No provider-dollar claim; unknown actual prices remain null.", + "cache_economic_guard": "Require upper paired bounds on weighted input and total candidate minus 1.01*C <=0 at all four sensitivity corners, over whole session and turns 1/2/3, with simultaneous correction specified in analysis.bootstrap. Check worst-case missing-usage intervals. No robust win if weights reverse ranking or an unbounded row could reverse it.", + "tie_rule": "Within 1% of minimum all-attempt tokens is a practical tie. Prefer C, then D0, D1, A0, A1 in that order; cost confidence intervals overlapping zero savings yield retain C pending another preregistration. No arbitrary latency tie-break after seeing data.", + "engine_rule": "Entire tested configurations only, confined to six named tasks/domains per role. This design cannot authorize moving a production builder, reviewer, researcher, prose-verdict or evidence-writing role. Reviewer/researcher style claims cover JSON-only records. Representative frozen tasks and graded prose records require a new preregistration.", + "incomplete": "Any unresolved sealing/qualification gate, missing matched outcome, resource stop or failed calibration yields incomplete/inconclusive and retain C. Keep all attempts and uncertainty; no filled-in pilot/resource values.", + "outcomes": { + "builder": null, + "reviewer": null, + "researcher": null + }, + "minimum_absolute_combined_success_rate": 0.9, + "absolute_quality_gate": "Control and candidate each require combined success >=0.90, with at least one real successful attempt of every selected task and nonempty exactness evidence. A trivially all-failing control cannot qualify a cheap treatment. A task omitted by the sealed random draw makes coverage incomplete; do not rerandomize after observations.", + "scope_ceiling": "named synthetic task domains only", + "review_research_style_scope": "JSON-only records" + }, + "role_policy": { + "early_canary_roles": [ + "builder" + ], + "control_until_pass": [ + "reviewer", + "researcher", + "judgment", + "verification", + "evidence" + ], + "record_outputs_require_exact_preservation": true, + "builder_canary": "Only isolated reversible builder trials after deterministic exactness and two-hop readiness; no early adoption, external publication or authoritative records.", + "trial_vs_production": "Reviewer/researcher treatment outputs are quarantined experimental artifacts graded independently on control/locally; their production routes remain 20128. No LLM grader in compressed arm.", + "promotion": "Role+configuration+task-domain specific; any broad review/evidence role needs representative confirmation and exact strings in verdict prose, receipts, pins and hashes. Stop on new record corruption and revert owned scope to control." + }, + "budget": { + "total_token_cap": null, + "allocations": null, + "unit": "Sum each arm's entry-gateway tokens_in + tokens_out once: C at 20128, D/A at 20129. Cache/reasoning subsets not added; exclude chained 20128 rows.", + "max_concurrent_sessions": 1, + "inflight_reserve_tokens": null, + "max_owned_session_wall_seconds": null, + "whole_run_wall_seconds": null, + "poll_seconds": 5, + "admission": "Before pilot freeze native per-request input/output caps and cancellation reserve. Before confirmation require full powered cohort fits sized allocations and available resource envelope. No optimistic partial run or reducing n to fit; preserve a failed feasibility result. With unenforced remote charge caps describe monitored soft budget plus overshoot bound, never a hard cap.", + "stop_conditions": [ + "Record or encrypted-reasoning corruption, failed negative controls, authentication/quota refusal, observed request-body model/effort/config drift or ambiguous correlation join.", + "Resource reserve or frozen wall limit reached; unknown unbounded inflight usage pauses admission without inventing zero usage.", + "Routine error/cancellation stop limits calibrated from pilot as specified by operational_stop_calibration; no arbitrary five-errors/three-of-ten rule.", + "Independent settings observation detects unauthorized management changes or uncontrolled lane traffic.", + "Missing qualified route, merged config, same native session, observer or network containment blocks start." + ], + "owned_process_cancellation": "Use native Harbor cancellation first; coordinator records own job/container/process group identities at launch. Stop launching, interrupt only that owned process group, allow 10 seconds, TERM then 10 seconds then KILL only surviving owned descendants; remove only owned sandbox resources through Harbor. Never pkill by model name, stop either gateway or kill unrelated sessions. Observe local death and separately reconcile each arm's entry-gateway terminal/499 rows for the pilot-frozen reconciliation window (at least 60 seconds). Remote provider cancellation unproven remains null; no automatic retry, no erased partial cohort.", + "restore": "Future host owner restores the exact owned 20129 experiment configuration snapshot; never changes 20128. No host state was configured by this draft.", + "actual_usage": null, + "sizing_rule": "After complete pilot choose powered n. Per role/cell use 1.25 times the one-sided 95% upper confidence limit on mean complete-session tokens/wall (including all retries/grading); total confirmation allocation is sum(n_role * cell upper mean). Also bootstrap the summed serial schedule at 99% and use the larger aggregate estimate. Add separately observed controls, actual pilot spending, a separately specified exploratory allocation (may be zero), and a reserve from qualified maximum inflight request/output/cancellation bounds. Serial wall allocation includes setup/teardown and sum of all five cells; set session cap at max(1800,2*pilot max seconds). Freeze values and uncertainty before confirmation. Statistical resource estimates are not hard charge bounds.", + "operational_stop_calibration": { + "thresholds": null, + "rule": "Before confirmation use pilot unrecovered-session and request-error counts to obtain one-sided 95% Clopper-Pearson upper rates. For each rolling 20-session block choose the smallest count k with binomial upper-tail probability <=.01/planned_blocks at that conservative rate (minimum allowance 2). Freeze poll delay timeout as max(30s,2*pilot p99 finalization delay) and reconciliation timeout similarly. Simulate temporal clustering as sensitivity; if no usable threshold exists or stationarity fails, remain unqualified. Authentication/config/corruption stops remain immediate." + } + }, + "threats_to_validity": [ + "Prompt cache: historical ~94% is not current warmth. The D/A live zone is keyed with TTL 60 but actual reuse, per-hop cache identity and per-turn cost remain unobserved.", + "Sharedgw adds a hop; default-only cells differ from C in hop and labelled-safe engines. No pure hop-only causal claim without additional globally-disabled/excluded framework control in a new protocol.", + "Session/account affinity, encrypted reasoning policy and fallback routing may change between hops; log actual identities privately and resolved model/effort.", + "Provider stochasticity and time-varying account capacity; fresh paired sessions, fixed order counterbalancing, same account affinity where supported, provider backend seed not assumed.", + "Five-cell time-sliced styles require exclusive ownership; config variants can disturb cache partition. Preserve snapshot hashes and report every transition.", + "Six fixed tasks per role and four builder smoke tasks constrain generalization; repetitions increase stochastic precision, not task diversity.", + "Promptfoo disk result cache would avoid model calls; use --no-cache. Keep native provider caching enabled and measured.", + "Enabled engine with no eligible input or missing optional model does not demonstrate that engine; engine interaction order and safe-label defects can dominate.", + "Repeated reads can be intentional; count only preregistered unchanged-file unrequested rereads; unknown artifact identity remains null." + ], + "evidence": { + "draft_class": "source-backed protocol and local structural validation only", + "upstream_tests_executed": false, + "model_sessions_executed": false, + "gateway_requests_executed": false, + "private_only": [ + "API keys", + "auth stores", + "raw native conversations", + "encrypted reasoning", + "account/session identities", + "personal host paths" + ], + "public": "Compact sanitized per-task/turn receipts, exact commands/exit results, upstream revision+file spans, hashes and equality booleans, failed attempts and unknown usage. Never generated summaries as upstream execution evidence." + }, + "sealing": { + "sealed_at": null, + "execution_revision": null, + "schedule_sha256": null, + "effective_configs_sha256": null, + "adapter_executables_sha256": null, + "image_digests": null, + "before_run": [ + "This repair is DRAFT, not frozen and not run. No execution is authorized here.", + "Before qualification, coordinator freezes qualification-only inputs/ceilings and closes route/config/network/privacy/controls gates needed to run it; live route probes belong to qualification, never confirmation.", + "Pilot data may determine n/resources only under the prespecified algorithm. Preserve raw pilot/failed conditions, independent observations and exclusions.", + "Before confirmation, close all applicable host gates, freeze n/resource/stop thresholds, exact binary/container/adapter/config/task hashes and realized schedule in a dated amendment. Unknowns remain null.", + "Material post-confirmatory changes require fresh cohort; no silent rewriting or pooling runners. Broader role/prose scope needs a separate preregistration, not a gate this cohort can pass." + ], + "amendments": "Append-only dated rationale plus old/new hashes and explicit before/after-data timing; never silently rewrite after observations. No run/adoption authority from this draft.", + "qualification_frozen_at": null, + "confirmation_frozen_at": null + }, + "anti_pattern_log": [ + { + "mistake": "Treating safe-default classification as proven losslessness.", + "correction": "Session-dedup index collision and lite truncation paths are visible; all cells require exactness checks.", + "verification": "or-engine-labels; or-dedup 291-345; or-lite 148-168; or-adapter 116-145; or-strategy 367-386." + }, + { + "mistake": "Treating explicit output styles and legacy caveman output mode as additive, or clearing only styles to disable them.", + "correction": "Explicit nonempty styles win; disable legacy fallback too.", + "verification": "or-styles 13-28; D0/A0 config contract." + }, + { + "mistake": "Assuming promptfoo --help is read-only.", + "correction": "Research preflight attempted logging/DB migration and failed with EROFS/SQLITE_CANTOPEN, exit 1; no successful host write observed. Use package metadata/source under no-host-change scope.", + "verification": "promptfoo-startup src/main.ts 63-64 before parse at 147; src/logger.ts 224-247." + }, + { + "mistake": "Assuming direct-shell gh failure means upstream research unavailable.", + "correction": "Installed research tool gh api succeeded; source byte comparisons and pin checks retained here.", + "verification": "Seventeen installed compression files compared byte-for-byte to OmniRoute a58000c7, all gh calls exit 0. Repair recomputed 17/17 installed hashes against retained pinned-source hashes." + }, + { + "mistake": "Assuming a configurable Harbor CODEX_HOME proves arbitrary parent-home or -p forwarding.", + "correction": "Superseded in repair: use native semantic merged-config profile route through Harbor config; require resolved-config and prompt-input equivalence, not literal -p forwarding.", + "verification": "harbor-codex 81,1352-1359,1432-1448; harbor-options 50-70." + }, + { + "mistake": "Leaving the peer-reported Responses-to-lite path untraced.", + "correction": "Installed bodyAdapter maps function_call_output to role tool; strategySelector invokes lite through it, and lite truncates eligible string content without consulting codex-responses metadata. This is source reachability, not observed live application.", + "verification": "or-adapter 116-145,205-221; or-strategy 367-386; or-lite 148-168,259-262." + }, + { + "mistake": "Applying a binomial gate to a fixed balanced heterogeneous task schedule.", + "correction": "Independent task draws remain, but a harmful-discordance-only bound has near-zero power for equal arms at .95 success. Use pilot-sized paired net difference and exact harmful-minus-beneficial fallback.", + "verification": "Repair recomputed n=240, alpha=.05/24: at most 3 harmful discordances, pass probability .003102 for identical independent .95 arms. Sources: scipy, scipy-exact; B5 structural red-first test." + }, + { + "mistake": "Guessing affinity-header names for the draft checklist.", + "correction": "Earlier correction was wrong: x-omniroute-connection pins a connection; session affinity comes from three session headers, body identifiers, prompt_cache_key, then input hash.", + "verification": "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/services/sessionAffinityPin.ts#L197-L221; tests.test_gpt6_lane_compression_ab_preregistration.test_sensitivity_affinity_and_scope_are_explicit failed before repair." + }, + { + "mistake": "Review B1: superseded preregistration assumption at f201e07b.", + "correction": "Canonical hop names fixed; retain Harbor option (b) with unresolved supported route, full prompt equivalence and live-200-max gates; compare three alternatives.", + "verification": "Verified stripping in v0.23.0/main. Established namespace/route facts reused. PLAUSIBLE bare-name 401 not re-probed; no working route asserted.", + "citations": [ + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1339-L1449", + "https://github.com/harbor-framework/harbor/blob/3c82380859d187957cfd5cd64802b076d9779550/src/harbor/agents/installed/codex.py#L1502-L1605", + "https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/providers/openai/codex-sdk.ts#L1034-L1134", + "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/sdk/typescript/src/exec.ts#L91-L178", + "https://github.com/meridianlabs-ai/inspect_swe/blob/7eb8dd64309db4cd0f6bdf1d0ffd9786a74a4088/src/inspect_swe/_codex_cli/codex_cli.py#L483-L637" + ] + }, + { + "mistake": "Review B2: superseded preregistration assumption at f201e07b.", + "correction": "Merged semantic config, actual forced command, equal-arm deviations, host-network overlay/probe, admin containment gate and native [[steps]] mapping.", + "verification": "Read loader/merge, Harbor upload/command/compose/multi-step code. WSL2 reachability and config equivalence unobserved, gated.", + "citations": [ + "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/config/src/loader/mod.rs#L286-L340", + "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/config/src/merge.rs#L56-L185", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1339-L1449", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/environments/docker/docker.py#L350-L420", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/trial/multi_step.py#L25-L110" + ] + }, + { + "mistake": "Review B3: superseded preregistration assumption at f201e07b.", + "correction": "Pre-confirmatory pilot measures all cell costs/wall times; powered n, confirmation budget/reserve/wall fields now null pending sizing. Zero primes.", + "verification": "Old arithmetic verified: 5,400 sessions, 3,703.70 tokens/session, 112 s/session. Historical 13,806-token row is a different invocation, not a measured lower bound for this cohort; no new host receipt read.", + "citations": [ + "https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_resampling.py#L300-L394", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/trial/multi_step.py#L25-L110" + ] + }, + { + "mistake": "Review B4: superseded preregistration assumption at f201e07b.", + "correction": "Usage at entry only joined on X-Correlation-Id; upstream-hook observer specified; no hop summation. Request-body effort at both hops; nullable columns corroborate only.", + "verification": "Verified call_logs schema/conditional effort and header emission; native Codex header export not found in inspected SSE path. Observer and second-hop exact join remain unqualified.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/handlers/chatHelpers.ts#L1172-L1184", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L628-L653", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/app/api/usage/call-logs/[id]/route.ts#L1-L22", + "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/codex-api/src/sse/responses.rs#L30-L101", + "https://github.com/mitmproxy/mitmproxy/blob/6c09d56e4c29a92f5ad01b03199977584b8ea14f/mitmproxy/proxy/layers/http/_hooks.py#L7-L39" + ] + }, + { + "mistake": "Review B5: superseded preregistration assumption at f201e07b.", + "correction": "Paired net-harm bootstrap with explicit conservative exact fallback; 12 candidate intersection-union Holm tests, unrecovered failures, pilot power and null calibration.", + "verification": "Recomputed reviewer's old-gate probabilities .003102/.306404/.784372; cited pinned bootstrap/exact/Holm code. No power result claimed without pilot.", + "citations": [ + "https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_resampling.py#L300-L394", + "https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_binomtest.py#L52-L115", + "https://github.com/statsmodels/statsmodels/blob/278ff9950636cdd4939b4055e339a8e681d79cab/statsmodels/stats/multitest.py#L99-L149", + "https://doi.org/10.1002/(SICI)1097-0258(19980430)17:8%3C891::AID-SIM780%3E3.0.CO;2-B", + "https://doi.org/10.1214/ss/1032280304", + "https://github.com/numpy/numpy/blob/c5ab79c14c98bfda1e60770ffa23a6130f8267b7/numpy/random/_generator.pyx#L299-L310" + ] + }, + { + "mistake": "Review M1: superseded preregistration assumption at f201e07b.", + "correction": "Registered-key D/A cells and TTL 60 are confirmatory; key name through env_key/extra_env templates only; reuse acceptance gated.", + "verification": "Established host configuration reused; inspected principal/session and environment-template source, no key file access.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/liveZone.ts#L127-L132", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/base.py#L560-L648", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/utils/env.py#L4-L65" + ] + }, + { + "mistake": "Review M2: superseded preregistration assumption at f201e07b.", + "correction": "Sensitivity cache ratio [0,1], output ratio [1,10]; robust all-corner/turn guard without claiming actual prices or subscription meter.", + "verification": "These ranges are explicit protocol assumptions over source-defined token components; actual price fields remain null.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107", + "https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_resampling.py#L300-L394" + ] + }, + { + "mistake": "Review M3: superseded preregistration assumption at f201e07b.", + "correction": "Four native output-shape canaries with whole-string eligibility and native-reachability results; shell minifier nonreachability cannot count as engine acceptance.", + "verification": "Verified shell/unified_exec headers and whole-string JSON.parse. MCP/custom shapes and actual engine application remain qualification results.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/codexResponses/index.ts#L137-L145", + "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/tools/context.rs#L524-L601", + "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/tools/mod.rs#L97-L124", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/bodyAdapter.ts#L116-L145" + ] + }, + { + "mistake": "Review M4: superseded preregistration assumption at f201e07b.", + "correction": "Direct Chat key-2 collision stimulus rebuilt; assert multipart pin part unchanged; native Responses reachability and wire format independently gated.", + "verification": "Source key construction/owner ordering/replacement verified. input_text bypass is source-supported concern; no native collision observed.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/session-dedup/index.ts#L291-L345", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/bodyAdapter.ts#L116-L145", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/strategySelector.ts#L367-L386" + ] + }, + { + "mistake": "Review M5: superseded preregistration assumption at f201e07b.", + "correction": "Reviewer/researcher style conclusions restricted to JSON-only records; no prose-verdict/evidence harmlessness claim.", + "verification": "All 12 task schemas inspected; style catalog targets prose/code. PLAUSIBLE effect magnitude remains unmeasured; narrowed scope uses reviewer's alternative fix.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/outputStyles/catalog.ts#L32-L194", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449" + ] + }, + { + "mistake": "Review M6: superseded preregistration assumption at f201e07b.", + "correction": "Connection pin separated from session affinity; prompt_cache_key at both hops, account spread/cache reads by arm, prior false correction repaired.", + "verification": "Header/body/fallback precedence verified. PLAUSIBLE account collapse not observed and not claimed; qualification required.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/services/sessionAffinityPin.ts#L197-L221", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/executors/codex.ts#L1498-L1573" + ] + }, + { + "mistake": "Review M7: superseded preregistration assumption at f201e07b.", + "correction": "Every session including C scores exact values; application/reachability is separate engine coverage, never exclusion or vacuous pass.", + "verification": "Source allows no-op/ineligible engine paths; public document contract now explicitly defines both outcomes.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/codexResponses/index.ts#L137-L145", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/strategySelector.ts#L367-L386" + ] + }, + { + "mistake": "Review M8: superseded preregistration assumption at f201e07b.", + "correction": "Drop all separate priming and cold/warm labels; stratify measured cache analysis by turn.", + "verification": "PLAUSIBLE lack of warming not established experimentally; native session/affinity source supports avoiding an unverified priming benefit. No cache-effect claim.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/services/sessionAffinityPin.ts#L197-L221", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/executors/codex.ts#L1498-L1573" + ] + }, + { + "mistake": "Review M9: superseded preregistration assumption at f201e07b.", + "correction": "Pilot-calibrated operational stops; null usage intervals and worst-case ranking. Request size alone is declined as a bound on total input+output usage.", + "verification": "Nullable input/output/reasoning schema verified. Routine error prevalence unmeasured. Require supported output/serialization caps as well as stored request size; without them cost gate stays open.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L628-L653", + "https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_binomtest.py#L52-L115" + ] + }, + { + "mistake": "Review M10: superseded preregistration assumption at f201e07b.", + "correction": "Explicit ceiling: named synthetic task domains only, no production role migration or prose-record adoption.", + "verification": "Six-task packets and four builder smoke cases inspected; narrower scope states the design's limitation without adding representative tasks.", + "citations": [ + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/verifier/verifier.py#L165-L250", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/outputStyles/catalog.ts#L32-L194" + ] + }, + { + "mistake": "Review m1: superseded preregistration assumption at f201e07b.", + "correction": "Replace Nine with Seventeen installed source files in correction log.", + "verification": "Recomputed all 17 retained installed SHA256 bindings: 17 matched. No new live acceptance implied.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/codexResponses/index.ts#L137-L145" + ] + }, + { + "mistake": "Review m2: superseded preregistration assumption at f201e07b.", + "correction": "Per-call tokens_compressed primary diagnostic; includes reactive compaction, not billed savings. Analytics optional, never additive.", + "verification": "Inspected callLogs L107 and chatCore L1916-1919/L2164; source-backed attribution improvement.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449" + ] + }, + { + "mistake": "Review m3: superseded preregistration assumption at f201e07b.", + "correction": "promptfoo gateway example uses sharedgw/gpt-6-astra-max; Codex SDK alternate preserves the same one-slash slug.", + "verification": "Verified provider model propagation and SDK --model code path; no promptfoo execution.", + "citations": [ + "https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/providers/openai/responses.ts#L1195-L1210", + "https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/providers/openai/codex-sdk.ts#L1034-L1134", + "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/sdk/typescript/src/exec.ts#L91-L178" + ] + }, + { + "mistake": "Review m4: superseded preregistration assumption at f201e07b.", + "correction": "Require ambient OPENAI_API_KEY/auth-copy switches unset and use OMNIROUTE_FW_API_KEY template; verify no literal/partial secret logs.", + "verification": "Harbor can inject OpenAI credentials and KEY-sensitive env handling verified. Presence checks and redaction runtime acceptance left to coordinator; no values opened.", + "citations": [ + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1339-L1449", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/base.py#L560-L648", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/utils/env.py#L4-L65" + ] + }, + { + "mistake": "Review m5: superseded preregistration assumption at f201e07b.", + "correction": "Responses->Responses sharedgw format specified; no Chat fallback; real encrypted/custom-item preservation gated.", + "verification": "Adapter and pre/post-translation stage source inspected; configured/actual node format still unobserved.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/bodyAdapter.ts#L116-L145", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/strategySelector.ts#L367-L386" + ] + }, + { + "mistake": "Review m6: superseded preregistration assumption at f201e07b.", + "correction": "Live-zone session fallback derived from request body/provider/connection is explicit and distinct from affinity extraction.", + "verification": "Inspected chatCore L1868-1881 and sessionManager L103-145.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/sessionManager.ts#L103-L145", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449" + ] + }, + { + "mistake": "Review m7: superseded preregistration assumption at f201e07b.", + "correction": "Tests now enforce canonical model hops and entry gateway ledger, plus all structurally checkable blockers.", + "verification": "Before document repair: requested unittest command exit 1, 10 tests, 12 failures, including old L61/L122 expectations.", + "citations": [ + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1339-L1449", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/handlers/chatHelpers.ts#L1172-L1184", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107" + ] + }, + { + "mistake": "Review m8: superseded preregistration assumption at f201e07b.", + "correction": "Exploratory repetition unit explicit; keyed exploratory cell removed; original two builder steps are turns 1/2 and record step is turn 3.", + "verification": "Pinned task.toml contains exactly create-file and append-content; native resume implementation inspected. Wrapper remains local integration, not unchanged upstream execution.", + "citations": [ + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/examples/tasks/hello-multi-step-simple/task.toml#L15-L31", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/trial/multi_step.py#L25-L110" + ] + }, + { + "mistake": "Assuming the requested alternate TMPDIR avoided all repository sentinels.", + "correction": "The required regression command found an existing read-only .git in /var/tmp/claude-431 too. Preserve the 11 export-refusal errors; do not remove metadata or weaken tests. Repository classification passes independently.", + "verification": "ls -ld showed dr-xr-xr-x for the sentinel; blind_checkout.py:970-973 directly checks ancestor/.git existence; 57 tests returned FAILED (errors=11).", + "enforced_by": "Existing blind export ancestor guard, unchanged." + }, + { + "mistake": "Describing tokens_compressed as selected compression plus reactive savings could imply addition.", + "correction": "The reactive path assigns tokensCompressed at chatCore.ts L2164, replacing the earlier selected-pipeline estimate. Label the row as latest recorded compression estimate; never sum stages.", + "verification": "Read installed chatCore.ts L1916-1919 and L2157-2164; source hash matches pinned upstream.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L2157-L2164" + ] + } + ], + "open_host_gates": [ + { + "id": "runner-route", + "requirement": "Supported 20129 slashless route preserving canonical native metadata, five-item/full prompt equivalence and coordinator live 200 at max; C route equivalence too. No mapping has been found/qualified.", + "passed": null + }, + { + "id": "merged-profile", + "requirement": "Materialize/hash native semantic base+stack-worker merge; no profile/profiles keys; same resolved config and prompt-input versus -p reference before/after Harbor forced flags.", + "passed": null + }, + { + "id": "network-containment", + "requirement": "Hash extra_docker_compose overlay; prove actual container reachability and separately qualify upstream-supported containment of passwordless admin APIs; trusted-task constraints alone are not a security boundary.", + "passed": null + }, + { + "id": "header-capture", + "requirement": "Materialize/hash upstream-hook response-header observer, qualify privacy/SSE/cache semantics and X-Correlation-Id entry joins including failures; demonstrate a separate exact second-hop effort join.", + "passed": null + }, + { + "id": "effort-detail", + "requirement": "Qualify detail-API selectors and inspect only reasoning.effort at joined hops; null corroboration columns are allowed, missing body stays unknown.", + "passed": null + }, + { + "id": "two-hop", + "requirement": "Freeze Responses wire format and prove three native steps, genuine call IDs/encrypted reasoning replay, prompt_cache_key, session/account affinity and arm account spread/cache rate.", + "passed": null + }, + { + "id": "effective-config", + "requirement": "Read back C off/exclusions; keyed D/A engine/style plans, 60-minute live zone, optional dependencies and safety compaction. Existing key/TTL are facts, reuse is not accepted.", + "passed": null + }, + { + "id": "native-harnesses", + "requirement": "Harbor 0.23.0, Codex 0.157.1, chosen observer, statistics, actual task dependencies/images and exact build/config hashes through supported upstream commands.", + "passed": null + }, + { + "id": "native-controls", + "requirement": "Materialize/hash three-turn wrappers and graders; native known-pass/fail/malformed controls, four numeric output shapes, rebuilt Chat collision plus separately observed native reachability.", + "passed": null + }, + { + "id": "pilot-power-resources", + "requirement": "Complete excluded pre-confirmatory pilot; calibrate NI at margin, simulate power and operational stops; freeze powered n, complete token budget, reserve, wall time and schedule.", + "passed": null + }, + { + "id": "usage-sensitivity", + "requirement": "Confirm usage fields and finalization, missing-row input/output bounds, all-corner sensitivity and simultaneous cost bounds; actual dollar prices remain unknown.", + "passed": null + } + ], + "local_validation": { + "classification": "local structural contract, not upstream acceptance", + "command": "rtk env PYTHONDONTWRITEBYTECODE=1 python3 -m unittest discover -s tests -p test_gpt6_lane_compression_ab_preregistration.py -v", + "fail_first": { + "exit_code": 1, + "result": "Ran 1 test; FAILED (failures=1): preregistration.json does not exist yet", + "before_draft_files": true + }, + "pass": { + "exit_code": 0, + "result": "Ran 1 test; OK", + "scope": "Local structural contract only; upstream/native/model execution not performed." + }, + "repository_validator": { + "command": "rtk env PYTHONDONTWRITEBYTECODE=1 GIT_OPTIONAL_LOCKS=0 python3 scripts/validate.py", + "exit_code": 0, + "result": { + "components": 69, + "hashed_files": 7348, + "profiles": 4, + "receipts": 159, + "status": "passed" + }, + "scope": "Repository integrity and scope checks only. Untracked draft files covered by their direct contract test and explicit publication scan." + }, + "publication_scan": { + "command": "rtk env PYTHONDONTWRITEBYTECODE=1 GIT_OPTIONAL_LOCKS=0 python3 scripts/validate.py --scan-file blueprints/gpt6-lane-compression-ab/PREREGISTRATION.md --scan-file blueprints/gpt6-lane-compression-ab/preregistration.json --scan-file tests/test_gpt6_lane_compression_ab_preregistration.py", + "exit_code": 0, + "result": { + "scanned_files": 3, + "status": "passed" + } + }, + "historical_scope": "Original authoring checks at/before f201e07b, not repair results; current checks are repair_round_20260927.checks." + }, + "qualification_pilot": { + "phase": "pre-confirmatory-data", + "excluded_from_confirmation": true, + "cells": [ + "C", + "D0", + "D1", + "A0", + "A1" + ], + "schedule": "12 complete paired draws per role, two of each six tasks, 180 sessions total; all five cells per draw, Williams order. Three turns per session, zero primes. Stop pilot on corruption/authorization/gate failure; never use an incomplete pilot to size confirmation. Diagnostic coverage pilot is deliberately balanced; confirmatory inference uses independent uniform draws.", + "measure": [ + "tokens", + "wall_seconds", + "paired_outcomes", + "operational_events", + "cache_by_turn", + "account_spread", + "null_usage_bounds", + "engine_reachability" + ], + "preliminary_ceiling": { + "tokens": 24000000, + "wall_seconds": 604800, + "session_wall_seconds": 1800, + "description": "Provisional pilot-only monitored ceiling, not a claim 180 sessions fit and not execution authority. Native per-request bounds/reserve qualification required before starting. If it cannot finish within ceiling, amend before confirmation; preserve partial pilot." + }, + "statistics": "Retain joint five-arm final success and unrecovered-failure indicators per draw and per role; retain task identities, costs, retries and timing. Compute per-cell mean/p95/max and upper confidence limits for complete-session all-attempt token/wall totals using pinned SciPy.", + "next_step": "Use pilot rates for analysis.power_simulation; derive complete-cohort n, cost reserve and wall allocation by budget.sizing_rule, then freeze confirmation schedule and hashes in a dated amendment before any confirmatory draw.", + "source_ids": [ + "harbor-multi-step", + "scipy", + "scipy-exact" + ], + "results": null + }, + "repair_round_20260927": { + "date": "2026-09-27", + "reviewed_head": "f201e07b", + "review_artifact": "431-review-claude.md", + "finding_count": 23, + "status": "DRAFT_not_frozen_not_run", + "pattern_source": "https://github.com/seathatflowsinourveins/native-agent-stack/blob/b6f36d8cac660b47cafc827ee4595cf0bc74459a/blueprints/compaction-window-ab/build-evidence.json", + "model_inference_performed": false, + "git_metadata_written": false, + "host_settings_changed": false, + "evidence_manifest_changed": false, + "subagents_spawned": 0, + "claude_sessions_started": 0, + "installed_hashes_checked": 17, + "installed_hashes_matched": 17, + "findings": [ + { + "id": "B1", + "outcome": "fixed by gating", + "reason": "Canonical hop names fixed; retain Harbor option (b) with unresolved supported route, full prompt equivalence and live-200-max gates; compare three alternatives.", + "verification": "Verified stripping in v0.23.0/main. Established namespace/route facts reused. PLAUSIBLE bare-name 401 not re-probed; no working route asserted.", + "citations": [ + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1339-L1449", + "https://github.com/harbor-framework/harbor/blob/3c82380859d187957cfd5cd64802b076d9779550/src/harbor/agents/installed/codex.py#L1502-L1605", + "https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/providers/openai/codex-sdk.ts#L1034-L1134", + "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/sdk/typescript/src/exec.ts#L91-L178", + "https://github.com/meridianlabs-ai/inspect_swe/blob/7eb8dd64309db4cd0f6bdf1d0ffd9786a74a4088/src/inspect_swe/_codex_cli/codex_cli.py#L483-L637" + ] + }, + { + "id": "B2", + "outcome": "fixed by gating", + "reason": "Merged semantic config, actual forced command, equal-arm deviations, host-network overlay/probe, admin containment gate and native [[steps]] mapping.", + "verification": "Read loader/merge, Harbor upload/command/compose/multi-step code. WSL2 reachability and config equivalence unobserved, gated.", + "citations": [ + "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/config/src/loader/mod.rs#L286-L340", + "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/config/src/merge.rs#L56-L185", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1339-L1449", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/environments/docker/docker.py#L350-L420", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/trial/multi_step.py#L25-L110" + ] + }, + { + "id": "B3", + "outcome": "fixed by gating", + "reason": "Pre-confirmatory pilot measures all cell costs/wall times; powered n, confirmation budget/reserve/wall fields now null pending sizing. Zero primes.", + "verification": "Old arithmetic verified: 5,400 sessions, 3,703.70 tokens/session, 112 s/session. Historical 13,806-token row is a different invocation, not a measured lower bound for this cohort; no new host receipt read.", + "citations": [ + "https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_resampling.py#L300-L394", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/trial/multi_step.py#L25-L110" + ] + }, + { + "id": "B4", + "outcome": "fixed by gating", + "reason": "Usage at entry only joined on X-Correlation-Id; upstream-hook observer specified; no hop summation. Request-body effort at both hops; nullable columns corroborate only.", + "verification": "Verified call_logs schema/conditional effort and header emission; native Codex header export not found in inspected SSE path. Observer and second-hop exact join remain unqualified.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/handlers/chatHelpers.ts#L1172-L1184", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L628-L653", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/app/api/usage/call-logs/[id]/route.ts#L1-L22", + "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/codex-api/src/sse/responses.rs#L30-L101", + "https://github.com/mitmproxy/mitmproxy/blob/6c09d56e4c29a92f5ad01b03199977584b8ea14f/mitmproxy/proxy/layers/http/_hooks.py#L7-L39" + ] + }, + { + "id": "B5", + "outcome": "fixed by gating", + "reason": "Paired net-harm bootstrap with explicit conservative exact fallback; 12 candidate intersection-union Holm tests, unrecovered failures, pilot power and null calibration.", + "verification": "Recomputed reviewer's old-gate probabilities .003102/.306404/.784372; cited pinned bootstrap/exact/Holm code. No power result claimed without pilot.", + "citations": [ + "https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_resampling.py#L300-L394", + "https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_binomtest.py#L52-L115", + "https://github.com/statsmodels/statsmodels/blob/278ff9950636cdd4939b4055e339a8e681d79cab/statsmodels/stats/multitest.py#L99-L149", + "https://doi.org/10.1002/(SICI)1097-0258(19980430)17:8%3C891::AID-SIM780%3E3.0.CO;2-B", + "https://doi.org/10.1214/ss/1032280304", + "https://github.com/numpy/numpy/blob/c5ab79c14c98bfda1e60770ffa23a6130f8267b7/numpy/random/_generator.pyx#L299-L310" + ] + }, + { + "id": "M1", + "outcome": "fixed by gating", + "reason": "Registered-key D/A cells and TTL 60 are confirmatory; key name through env_key/extra_env templates only; reuse acceptance gated.", + "verification": "Established host configuration reused; inspected principal/session and environment-template source, no key file access.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/liveZone.ts#L127-L132", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/base.py#L560-L648", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/utils/env.py#L4-L65" + ] + }, + { + "id": "M2", + "outcome": "fixed", + "reason": "Sensitivity cache ratio [0,1], output ratio [1,10]; robust all-corner/turn guard without claiming actual prices or subscription meter.", + "verification": "These ranges are explicit protocol assumptions over source-defined token components; actual price fields remain null.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107", + "https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_resampling.py#L300-L394" + ] + }, + { + "id": "M3", + "outcome": "fixed by gating", + "reason": "Four native output-shape canaries with whole-string eligibility and native-reachability results; shell minifier nonreachability cannot count as engine acceptance.", + "verification": "Verified shell/unified_exec headers and whole-string JSON.parse. MCP/custom shapes and actual engine application remain qualification results.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/codexResponses/index.ts#L137-L145", + "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/tools/context.rs#L524-L601", + "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/codex-rs/core/src/tools/mod.rs#L97-L124", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/bodyAdapter.ts#L116-L145" + ] + }, + { + "id": "M4", + "outcome": "fixed by gating", + "reason": "Direct Chat key-2 collision stimulus rebuilt; assert multipart pin part unchanged; native Responses reachability and wire format independently gated.", + "verification": "Source key construction/owner ordering/replacement verified. input_text bypass is source-supported concern; no native collision observed.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/session-dedup/index.ts#L291-L345", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/bodyAdapter.ts#L116-L145", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/strategySelector.ts#L367-L386" + ] + }, + { + "id": "M5", + "outcome": "fixed", + "reason": "Reviewer/researcher style conclusions restricted to JSON-only records; no prose-verdict/evidence harmlessness claim.", + "verification": "All 12 task schemas inspected; style catalog targets prose/code. PLAUSIBLE effect magnitude remains unmeasured; narrowed scope uses reviewer's alternative fix.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/outputStyles/catalog.ts#L32-L194", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449" + ] + }, + { + "id": "M6", + "outcome": "fixed by gating", + "reason": "Connection pin separated from session affinity; prompt_cache_key at both hops, account spread/cache reads by arm, prior false correction repaired.", + "verification": "Header/body/fallback precedence verified. PLAUSIBLE account collapse not observed and not claimed; qualification required.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/services/sessionAffinityPin.ts#L197-L221", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/executors/codex.ts#L1498-L1573" + ] + }, + { + "id": "M7", + "outcome": "fixed", + "reason": "Every session including C scores exact values; application/reachability is separate engine coverage, never exclusion or vacuous pass.", + "verification": "Source allows no-op/ineligible engine paths; public document contract now explicitly defines both outcomes.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/codexResponses/index.ts#L137-L145", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/strategySelector.ts#L367-L386" + ] + }, + { + "id": "M8", + "outcome": "fixed", + "reason": "Drop all separate priming and cold/warm labels; stratify measured cache analysis by turn.", + "verification": "PLAUSIBLE lack of warming not established experimentally; native session/affinity source supports avoiding an unverified priming benefit. No cache-effect claim.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/services/sessionAffinityPin.ts#L197-L221", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/executors/codex.ts#L1498-L1573" + ] + }, + { + "id": "M9", + "outcome": "fixed by gating", + "reason": "Pilot-calibrated operational stops; null usage intervals and worst-case ranking. Request size alone is declined as a bound on total input+output usage.", + "verification": "Nullable input/output/reasoning schema verified. Routine error prevalence unmeasured. Require supported output/serialization caps as well as stored request size; without them cost gate stays open.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L628-L653", + "https://github.com/scipy/scipy/blob/e4e854eaa8f18d807cd3496028e257e36caa93cc/scipy/stats/_binomtest.py#L52-L115" + ] + }, + { + "id": "M10", + "outcome": "fixed", + "reason": "Explicit ceiling: named synthetic task domains only, no production role migration or prose-record adoption.", + "verification": "Six-task packets and four builder smoke cases inspected; narrower scope states the design's limitation without adding representative tasks.", + "citations": [ + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/verifier/verifier.py#L165-L250", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/outputStyles/catalog.ts#L32-L194" + ] + }, + { + "id": "m1", + "outcome": "fixed", + "reason": "Replace Nine with Seventeen installed source files in correction log.", + "verification": "Recomputed all 17 retained installed SHA256 bindings: 17 matched. No new live acceptance implied.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/engines/codexResponses/index.ts#L137-L145" + ] + }, + { + "id": "m2", + "outcome": "fixed", + "reason": "Per-call tokens_compressed primary diagnostic; includes reactive compaction, not billed savings. Analytics optional, never additive.", + "verification": "Inspected callLogs L107 and chatCore L1916-1919/L2164; source-backed attribution improvement.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449" + ] + }, + { + "id": "m3", + "outcome": "fixed", + "reason": "promptfoo gateway example uses sharedgw/gpt-6-astra-max; Codex SDK alternate preserves the same one-slash slug.", + "verification": "Verified provider model propagation and SDK --model code path; no promptfoo execution.", + "citations": [ + "https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/providers/openai/responses.ts#L1195-L1210", + "https://github.com/promptfoo/promptfoo/blob/34f74d34e140b5e17d23770dfb2340057b1936b8/src/providers/openai/codex-sdk.ts#L1034-L1134", + "https://github.com/openai/codex/blob/36650394c5b38c2990ccf2a3457165ca3e9d9726/sdk/typescript/src/exec.ts#L91-L178" + ] + }, + { + "id": "m4", + "outcome": "fixed by gating", + "reason": "Require ambient OPENAI_API_KEY/auth-copy switches unset and use OMNIROUTE_FW_API_KEY template; verify no literal/partial secret logs.", + "verification": "Harbor can inject OpenAI credentials and KEY-sensitive env handling verified. Presence checks and redaction runtime acceptance left to coordinator; no values opened.", + "citations": [ + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1339-L1449", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/base.py#L560-L648", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/utils/env.py#L4-L65" + ] + }, + { + "id": "m5", + "outcome": "fixed by gating", + "reason": "Responses->Responses sharedgw format specified; no Chat fallback; real encrypted/custom-item preservation gated.", + "verification": "Adapter and pre/post-translation stage source inspected; configured/actual node format still unobserved.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/bodyAdapter.ts#L116-L145", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/compression/strategySelector.ts#L367-L386" + ] + }, + { + "id": "m6", + "outcome": "fixed", + "reason": "Live-zone session fallback derived from request body/provider/connection is explicit and distinct from affinity extraction.", + "verification": "Inspected chatCore L1868-1881 and sessionManager L103-145.", + "citations": [ + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/services/sessionManager.ts#L103-L145", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/open-sse/handlers/chatCore.ts#L1425-L1449" + ] + }, + { + "id": "m7", + "outcome": "fixed", + "reason": "Tests now enforce canonical model hops and entry gateway ledger, plus all structurally checkable blockers.", + "verification": "Before document repair: requested unittest command exit 1, 10 tests, 12 failures, including old L61/L122 expectations.", + "citations": [ + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/agents/installed/codex.py#L1339-L1449", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/sse/handlers/chatHelpers.ts#L1172-L1184", + "https://github.com/diegosouzapw/OmniRoute/blob/a58000c7685f4091c7a6fd8ddf3ebce7d2ec67c3/src/lib/usage/callLogs.ts#L90-L107" + ] + }, + { + "id": "m8", + "outcome": "fixed", + "reason": "Exploratory repetition unit explicit; keyed exploratory cell removed; original two builder steps are turns 1/2 and record step is turn 3.", + "verification": "Pinned task.toml contains exactly create-file and append-content; native resume implementation inspected. Wrapper remains local integration, not unchanged upstream execution.", + "citations": [ + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/examples/tasks/hello-multi-step-simple/task.toml#L15-L31", + "https://github.com/harbor-framework/harbor/blob/1e5c5c6db929a10a140d05e606882c671ae20729/src/harbor/trial/multi_step.py#L25-L110" + ] + } + ], + "boundaries": "Only assigned worktree edited; temporary directory created as requested. PYTHONDONTWRITEBYTECODE=1; no key values read or copied. No model probes, host services, installs or external messages. No git add/commit.", + "research_failures": [ + "Direct shell gh API: exit 1 error connecting to api.github.com. Same read-only API through Context Mode succeeded.", + "Web open unsupported (HTTP 400); web search produced no usable source output. Pinned gh API sources used instead.", + "Sibling build-evidence file absent in this checkout; fetched exact PR #416 head b6f36d8c via gh API.", + "Speculative source paths examples/tasks/multi-step/task.toml and examples/addons/http-events.py returned 404; tree lookup located hello-multi-step-simple/task.toml and upstream HTTP hook/stream examples.", + "Installed API route source absent; pinned upstream src/app/api/usage/call-logs/[id]/route.ts inspected. Do not confuse an absent packaged source file with an absent API." + ], + "checks": [ + { + "phase": "repair_red_before_protocol_edits", + "command": "rtk env PYTHONDONTWRITEBYTECODE=1 TMPDIR=/var/tmp/claude-431 GIT_OPTIONAL_LOCKS=0 python3 -m unittest tests.test_gpt6_lane_compression_ab_preregistration -v", + "exit_code": 1, + "returned_output": "test_all_review_findings_have_cited_round_dispositions (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_all_review_findings_have_cited_round_dispositions) ... FAIL\ntest_b1_per_hop_names_and_unqualified_runner_cannot_seal (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b1_per_hop_names_and_unqualified_runner_cannot_seal) ... FAIL\ntest_b2_merge_launch_network_and_native_resume (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b2_merge_launch_network_and_native_resume) ... FAIL\ntest_b3_pilot_sizes_complete_cohort_and_removes_priming (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b3_pilot_sizes_complete_cohort_and_removes_priming) ... FAIL\ntest_b4_entry_joins_effort_body_and_unknown_usage (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b4_entry_joins_effort_body_and_unknown_usage) ... FAIL\ntest_b5_net_difference_power_and_unrecovered_failure (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b5_net_difference_power_and_unrecovered_failure) ... FAIL\ntest_canaries_have_reachable_shape_gates_and_valid_collision (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_canaries_have_reachable_shape_gates_and_valid_collision) ... FAIL\ntest_draft_contract (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_draft_contract) ... \n test_draft_contract (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_draft_contract) (contract='clean control and crossed factors') ... FAIL\n test_draft_contract (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_draft_contract) (contract='measurement boundaries') ... FAIL\n test_draft_contract (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_draft_contract) (contract='statistics and stopping are specified') ... FAIL\ntest_keyed_confirmatory_env_and_no_ambient_openai_key (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_keyed_confirmatory_env_and_no_ambient_openai_key) ... FAIL\ntest_sensitivity_affinity_and_scope_are_explicit (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_sensitivity_affinity_and_scope_are_explicit) ... FAIL\n\n======================================================================\nFAIL: test_all_review_findings_have_cited_round_dispositions (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_all_review_findings_have_cited_round_dispositions)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_gpt6_lane_compression_ab_preregistration.py\", line 339, in test_all_review_findings_have_cited_round_dispositions\n self.assertEqual({r[\"id\"] for r in rows},\n ~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^\n {f\"B{i}\" for i in range(1, 6)} |\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n {f\"M{i}\" for i in range(1, 11)} |\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n {f\"m{i}\" for i in range(1, 9)})\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\nAssertionError: Items in the second set but not the first:\n'M5'\n'm2'\n'm8'\n'M1'\n'M7'\n'B3'\n'M9'\n'B1'\n'M10'\n'm5'\n'M2'\n'm4'\n'B4'\n'm3'\n'M8'\n'm1'\n'B2'\n'm7'\n'm6'\n'B5'\n'M3'\n'M6'\n'M4'\n\n======================================================================\nFAIL: test_b1_per_hop_names_and_unqualified_runner_cannot_seal (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b1_per_hop_names_and_unqualified_runner_cannot_seal)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_gpt6_lane_compression_ab_preregistration.py\", line 199, in test_b1_per_hop_names_and_unqualified_runner_cannot_seal\n self.assertEqual(data[\"client\"].get(\"model_hops\"), {\n ~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"client_to_20129\": \"sharedgw/gpt-6-astra-max\",\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"20129_to_20128\": \"gpt-6-astra-max\",\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"control_C\": \"cx/gpt-6-astra-max\",\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n })\n ^^\nAssertionError: None != {'client_to_20129': 'sharedgw/gpt-6-astra[73 chars]max'}\n\n======================================================================\nFAIL: test_b2_merge_launch_network_and_native_resume (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b2_merge_launch_network_and_native_resume)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_gpt6_lane_compression_ab_preregistration.py\", line 219, in test_b2_merge_launch_network_and_native_resume\n self.assertEqual(merge.get(\"inputs\"), [\"config.toml\", \"stack-worker.config.toml\"])\n ~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\nAssertionError: None != ['config.toml', 'stack-worker.config.toml']\n\n======================================================================\nFAIL: test_b3_pilot_sizes_complete_cohort_and_removes_priming (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b3_pilot_sizes_complete_cohort_and_removes_priming)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_gpt6_lane_compression_ab_preregistration.py\", line 246, in test_b3_pilot_sizes_complete_cohort_and_removes_priming\n self.assertEqual(pilot.get(\"phase\"), \"pre-confirmatory-data\")\n ~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\nAssertionError: None != 'pre-confirmatory-data'\n\n======================================================================\nFAIL: test_b4_entry_joins_effort_body_and_unknown_usage (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b4_entry_joins_effort_body_and_unknown_usage)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_gpt6_lane_compression_ab_preregistration.py\", line 259, in test_b4_entry_joins_effort_body_and_unknown_usage\n self.assertIn(\"correlation_id\", accounting[\"required_columns\"])\n ~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\nAssertionError: 'correlation_id' not found in ['timestamp', 'path', 'status', 'model', 'reasoning_effort_requested', 'reasoning_effort_upstream', 'tokens_in', 'tokens_cache_read', 'tokens_reasoning']\n\n======================================================================\nFAIL: test_b5_net_difference_power_and_unrecovered_failure (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b5_net_difference_power_and_unrecovered_failure)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_gpt6_lane_compression_ab_preregistration.py\", line 276, in test_b5_net_difference_power_and_unrecovered_failure\n self.assertEqual(ni.get(\"method\"), \"scipy.stats.bootstrap\")\n ~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\nAssertionError: None != 'scipy.stats.bootstrap'\n\n======================================================================\nFAIL: test_canaries_have_reachable_shape_gates_and_valid_collision (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_canaries_have_reachable_shape_gates_and_valid_collision)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_gpt6_lane_compression_ab_preregistration.py\", line 306, in test_canaries_have_reachable_shape_gates_and_valid_collision\n self.assertEqual({row[\"shape\"] for row in packet.get(\"numeric_shapes\", [])},\n ~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n {\"shell\", \"mcp-text\", \"mcp-structuredContent\", \"custom-tool\"})\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\nAssertionError: Items in the second set but not the first:\n'custom-tool'\n'mcp-text'\n'shell'\n'mcp-structuredContent'\n\n======================================================================\nFAIL: test_draft_contract (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_draft_contract) (contract='clean control and crossed factors')\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_gpt6_lane_compression_ab_preregistration.py\", line 61, in test_draft_contract\n self.assertEqual(cell[\"model\"], \"sharedgw/gpt-6-astra-max\")\n ~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\nAssertionError: 'sharedgw/cx/gpt-6-astra-max' != 'sharedgw/gpt-6-astra-max'\n- sharedgw/cx/gpt-6-astra-max\n? ---\n+ sharedgw/gpt-6-astra-max\n\n\n======================================================================\nFAIL: test_draft_contract (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_draft_contract) (contract='measurement boundaries')\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_gpt6_lane_compression_ab_preregistration.py\", line 122, in test_draft_contract\n self.assertEqual(accounting[\"usage_authority\"], {\n ~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"C\": \"20128.call_logs\", \"D0\": \"20129.call_logs\",\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"D1\": \"20129.call_logs\", \"A0\": \"20129.call_logs\",\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"A1\": \"20129.call_logs\",\n ^^^^^^^^^^^^^^^^^^^^^^^^\n })\n ^^\nAssertionError: '20128.call_logs' != {'C': '20128.call_logs', 'D0': '20129.cal[78 chars]ogs'}\n\n======================================================================\nFAIL: test_draft_contract (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_draft_contract) (contract='statistics and stopping are specified')\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_gpt6_lane_compression_ab_preregistration.py\", line 154, in test_draft_contract\n self.assertIsNone(analysis[\"planned_repetitions_per_role_cell\"])\n ~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\nAssertionError: 240 is not None\n\n======================================================================\nFAIL: test_keyed_confirmatory_env_and_no_ambient_openai_key (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_keyed_confirmatory_env_and_no_ambient_openai_key)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_gpt6_lane_compression_ab_preregistration.py\", line 295, in test_keyed_confirmatory_env_and_no_ambient_openai_key\n self.assertEqual(cell[\"principal\"], \"registered-lane-key\")\n ~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\nAssertionError: 'keyless' != 'registered-lane-key'\n- keyless\n+ registered-lane-key\n\n\n======================================================================\nFAIL: test_sensitivity_affinity_and_scope_are_explicit (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_sensitivity_affinity_and_scope_are_explicit)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_gpt6_lane_compression_ab_preregistration.py\", line 324, in test_sensitivity_affinity_and_scope_are_explicit\n self.assertEqual(ranges.get(\"cache_ratio\"), [0.0, 1.0])\n ~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\nAssertionError: None != [0.0, 1.0]\n\n----------------------------------------------------------------------\nRan 10 tests in 0.009s\n\nFAILED (failures=12)\n", + "evidence_class": "local structural/integrity validation" + }, + { + "phase": "repair_green_after_protocol_edits", + "command": "rtk env PYTHONDONTWRITEBYTECODE=1 TMPDIR=/var/tmp/claude-431 GIT_OPTIONAL_LOCKS=0 python3 -m unittest tests.test_gpt6_lane_compression_ab_preregistration -v", + "exit_code": 0, + "returned_output": "test_all_review_findings_have_cited_round_dispositions (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_all_review_findings_have_cited_round_dispositions) ... ok\ntest_b1_per_hop_names_and_unqualified_runner_cannot_seal (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b1_per_hop_names_and_unqualified_runner_cannot_seal) ... ok\ntest_b2_merge_launch_network_and_native_resume (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b2_merge_launch_network_and_native_resume) ... ok\ntest_b3_pilot_sizes_complete_cohort_and_removes_priming (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b3_pilot_sizes_complete_cohort_and_removes_priming) ... ok\ntest_b4_entry_joins_effort_body_and_unknown_usage (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b4_entry_joins_effort_body_and_unknown_usage) ... ok\ntest_b5_net_difference_power_and_unrecovered_failure (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b5_net_difference_power_and_unrecovered_failure) ... ok\ntest_canaries_have_reachable_shape_gates_and_valid_collision (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_canaries_have_reachable_shape_gates_and_valid_collision) ... ok\ntest_draft_contract (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_draft_contract) ... ok\ntest_keyed_confirmatory_env_and_no_ambient_openai_key (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_keyed_confirmatory_env_and_no_ambient_openai_key) ... ok\ntest_sensitivity_affinity_and_scope_are_explicit (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_sensitivity_affinity_and_scope_are_explicit) ... ok\n\n----------------------------------------------------------------------\nRan 10 tests in 0.011s\n\nOK\n", + "evidence_class": "local structural/integrity validation" + }, + { + "phase": "required_regression_suite", + "command": "rtk env PYTHONDONTWRITEBYTECODE=1 TMPDIR=/var/tmp/claude-431 GIT_OPTIONAL_LOCKS=0 python3 -m unittest tests.test_osv_lockfile_coverage tests.test_blind_checkout tests.test_workflow_security_coverage", + "exit_code": 1, + "returned_output": "..........EEEEEE.E.EE....EE..............................\n======================================================================\nERROR: test_a_reference_to_a_symlink_whose_target_is_not_exported_is_missing (tests.test_blind_checkout.AllowlistExportTests.test_a_reference_to_a_symlink_whose_target_is_not_exported_is_missing)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_blind_checkout.py\", line 670, in setUp\n self.assertEqual(blind_checkout.main([\"--source\", str(self.source), \"--rev\", \"HEAD\",\n ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--dest\", str(self.dest), \"--export\", str(self.export),\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--allow-from-packets\", str(self.packets),\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--allow-missing-refs\"]), 0)\n ^^^^^^^^^^^^^^^^^^^^^^^^\n File \"/tools/sota-convergence/blind_checkout.py\", line 972, in main\n raise SystemExit(f\"--export {export} is inside the git repository {ancestor}; place it outside every \"\n \"repository\")\nSystemExit: --export /hosts/blind/export is inside the git repository /var/tmp/claude-431; place it outside every repository\n\n======================================================================\nERROR: test_an_export_inside_the_worktree_is_refused (tests.test_blind_checkout.AllowlistExportTests.test_an_export_inside_the_worktree_is_refused)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_blind_checkout.py\", line 670, in setUp\n self.assertEqual(blind_checkout.main([\"--source\", str(self.source), \"--rev\", \"HEAD\",\n ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--dest\", str(self.dest), \"--export\", str(self.export),\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--allow-from-packets\", str(self.packets),\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--allow-missing-refs\"]), 0)\n ^^^^^^^^^^^^^^^^^^^^^^^^\n File \"/tools/sota-convergence/blind_checkout.py\", line 972, in main\n raise SystemExit(f\"--export {export} is inside the git repository {ancestor}; place it outside every \"\n \"repository\")\nSystemExit: --export /hosts/blind/export is inside the git repository /var/tmp/claude-431; place it outside every repository\n\n======================================================================\nERROR: test_bare_reference_reduction (tests.test_blind_checkout.AllowlistExportTests.test_bare_reference_reduction)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_blind_checkout.py\", line 670, in setUp\n self.assertEqual(blind_checkout.main([\"--source\", str(self.source), \"--rev\", \"HEAD\",\n ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--dest\", str(self.dest), \"--export\", str(self.export),\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--allow-from-packets\", str(self.packets),\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--allow-missing-refs\"]), 0)\n ^^^^^^^^^^^^^^^^^^^^^^^^\n File \"/tools/sota-convergence/blind_checkout.py\", line 972, in main\n raise SystemExit(f\"--export {export} is inside the git repository {ancestor}; place it outside every \"\n \"repository\")\nSystemExit: --export /hosts/blind/export is inside the git repository /var/tmp/claude-431; place it outside every repository\n\n======================================================================\nERROR: test_counts_and_missing_references_are_reported (tests.test_blind_checkout.AllowlistExportTests.test_counts_and_missing_references_are_reported)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_blind_checkout.py\", line 670, in setUp\n self.assertEqual(blind_checkout.main([\"--source\", str(self.source), \"--rev\", \"HEAD\",\n ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--dest\", str(self.dest), \"--export\", str(self.export),\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--allow-from-packets\", str(self.packets),\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--allow-missing-refs\"]), 0)\n ^^^^^^^^^^^^^^^^^^^^^^^^\n File \"/tools/sota-convergence/blind_checkout.py\", line 972, in main\n raise SystemExit(f\"--export {export} is inside the git repository {ancestor}; place it outside every \"\n \"repository\")\nSystemExit: --export /hosts/blind/export is inside the git repository /var/tmp/claude-431; place it outside every repository\n\n======================================================================\nERROR: test_export_holds_only_referenced_and_transitive_paths (tests.test_blind_checkout.AllowlistExportTests.test_export_holds_only_referenced_and_transitive_paths)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_blind_checkout.py\", line 670, in setUp\n self.assertEqual(blind_checkout.main([\"--source\", str(self.source), \"--rev\", \"HEAD\",\n ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--dest\", str(self.dest), \"--export\", str(self.export),\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--allow-from-packets\", str(self.packets),\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--allow-missing-refs\"]), 0)\n ^^^^^^^^^^^^^^^^^^^^^^^^\n File \"/tools/sota-convergence/blind_checkout.py\", line 972, in main\n raise SystemExit(f\"--export {export} is inside the git repository {ancestor}; place it outside every \"\n \"repository\")\nSystemExit: --export /hosts/blind/export is inside the git repository /var/tmp/claude-431; place it outside every repository\n\n======================================================================\nERROR: test_missing_references_are_refused_by_default (tests.test_blind_checkout.AllowlistExportTests.test_missing_references_are_refused_by_default)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_blind_checkout.py\", line 670, in setUp\n self.assertEqual(blind_checkout.main([\"--source\", str(self.source), \"--rev\", \"HEAD\",\n ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--dest\", str(self.dest), \"--export\", str(self.export),\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--allow-from-packets\", str(self.packets),\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--allow-missing-refs\"]), 0)\n ^^^^^^^^^^^^^^^^^^^^^^^^\n File \"/tools/sota-convergence/blind_checkout.py\", line 972, in main\n raise SystemExit(f\"--export {export} is inside the git repository {ancestor}; place it outside every \"\n \"repository\")\nSystemExit: --export /hosts/blind/export is inside the git repository /var/tmp/claude-431; place it outside every repository\n\n======================================================================\nERROR: test_no_catalog_json_in_the_export_carries_a_winner_or_incumbent_value (tests.test_blind_checkout.CatalogWinnerKeyTests.test_no_catalog_json_in_the_export_carries_a_winner_or_incumbent_value)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_blind_checkout.py\", line 567, in test_no_catalog_json_in_the_export_carries_a_winner_or_incumbent_value\n self.assertEqual(blind_checkout.main([\"--source\", str(self.source), \"--rev\", \"HEAD\",\n ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--dest\", str(self.dest), \"--export\", str(export)]), 0)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File \"/tools/sota-convergence/blind_checkout.py\", line 972, in main\n raise SystemExit(f\"--export {export} is inside the git repository {ancestor}; place it outside every \"\n \"repository\")\nSystemExit: --export /hosts/blind/export is inside the git repository /var/tmp/claude-431; place it outside every repository\n\n======================================================================\nERROR: test_escaping_symlinks_are_removed_and_inner_ones_kept (tests.test_blind_checkout.ExportSymlinkTests.test_escaping_symlinks_are_removed_and_inner_ones_kept)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_blind_checkout.py\", line 482, in test_escaping_symlinks_are_removed_and_inner_ones_kept\n self.assertEqual(blind_checkout.main([\"--source\", str(self.source), \"--rev\", \"HEAD\",\n ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--dest\", str(self.dest), \"--export\", str(export)]), 0)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File \"/tools/sota-convergence/blind_checkout.py\", line 972, in main\n raise SystemExit(f\"--export {export} is inside the git repository {ancestor}; place it outside every \"\n \"repository\")\nSystemExit: --export /hosts/blind/export is inside the git repository /var/tmp/claude-431; place it outside every repository\n\n======================================================================\nERROR: test_instruction_names_that_are_directory_symlinks_become_stubs (tests.test_blind_checkout.ExportSymlinkTests.test_instruction_names_that_are_directory_symlinks_become_stubs)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_blind_checkout.py\", line 504, in test_instruction_names_that_are_directory_symlinks_become_stubs\n self.assertEqual(blind_checkout.main([\"--source\", str(self.source), \"--rev\", \"HEAD\",\n ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--dest\", str(self.dest), \"--export\", str(export)]), 0)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File \"/tools/sota-convergence/blind_checkout.py\", line 972, in main\n raise SystemExit(f\"--export {export} is inside the git repository {ancestor}; place it outside every \"\n \"repository\")\nSystemExit: --export /hosts/blind/export is inside the git repository /var/tmp/claude-431; place it outside every repository\n\n======================================================================\nERROR: test_export_has_the_stripped_tree_without_git_or_the_manifest (tests.test_blind_checkout.ExportTests.test_export_has_the_stripped_tree_without_git_or_the_manifest)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_blind_checkout.py\", line 386, in test_export_has_the_stripped_tree_without_git_or_the_manifest\n self.assertEqual(blind_checkout.main([\"--source\", str(self.source), \"--rev\", \"HEAD\",\n ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--dest\", str(self.dest), \"--export\", str(export)]), 0)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File \"/tools/sota-convergence/blind_checkout.py\", line 972, in main\n raise SystemExit(f\"--export {export} is inside the git repository {ancestor}; place it outside every \"\n \"repository\")\nSystemExit: --export /hosts/blind/export is inside the git repository /var/tmp/claude-431; place it outside every repository\n\n======================================================================\nERROR: test_the_export_replaces_every_project_instruction_file_and_leaves_the_worktree_alone (tests.test_blind_checkout.ExportTests.test_the_export_replaces_every_project_instruction_file_and_leaves_the_worktree_alone)\nPR #141 review (P1): AGENTS.md and CLAUDE.md were copied unchanged, so a Codex child started with\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/tests/test_blind_checkout.py\", line 440, in test_the_export_replaces_every_project_instruction_file_and_leaves_the_worktree_alone\n self.assertEqual(blind_checkout.main([\"--source\", str(self.source), \"--rev\", \"HEAD\",\n ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \"--dest\", str(self.dest), \"--export\", str(export)]), 0)\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n File \"/tools/sota-convergence/blind_checkout.py\", line 972, in main\n raise SystemExit(f\"--export {export} is inside the git repository {ancestor}; place it outside every \"\n \"repository\")\nSystemExit: --export /hosts/blind/export is inside the git repository /var/tmp/claude-431; place it outside every repository\n\n----------------------------------------------------------------------\nRan 57 tests in 24.869s\n\nFAILED (errors=11)\n", + "evidence_class": "local structural/integrity validation", + "result": "57 tests; 11 errors all from an existing read-only .git in requested TMPDIR. No classification failure. No .git deletion, fixture bypass, test modification or host change to force a pass." + }, + { + "phase": "repository_validator", + "command": "rtk env PYTHONDONTWRITEBYTECODE=1 TMPDIR=/var/tmp/claude-431 GIT_OPTIONAL_LOCKS=0 python3 scripts/validate.py", + "exit_code": 0, + "returned_output": "{\"components\": 69, \"hashed_files\": 7348, \"profiles\": 4, \"receipts\": 159, \"status\": \"passed\"}\nIntegrity and scope checks only; no live provider or GPU execution.\n", + "evidence_class": "local structural/integrity validation" + }, + { + "phase": "whitespace_check", + "command": "rtk git diff --check", + "exit_code": 0, + "returned_output": "", + "evidence_class": "local structural/integrity validation" + }, + { + "phase": "post_record_contract", + "command": "rtk env PYTHONDONTWRITEBYTECODE=1 TMPDIR=/var/tmp/claude-431 GIT_OPTIONAL_LOCKS=0 python3 -m unittest tests.test_gpt6_lane_compression_ab_preregistration -v", + "exit_code": 0, + "returned_output": "test_all_review_findings_have_cited_round_dispositions (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_all_review_findings_have_cited_round_dispositions) ... ok\ntest_b1_per_hop_names_and_unqualified_runner_cannot_seal (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b1_per_hop_names_and_unqualified_runner_cannot_seal) ... ok\ntest_b2_merge_launch_network_and_native_resume (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b2_merge_launch_network_and_native_resume) ... ok\ntest_b3_pilot_sizes_complete_cohort_and_removes_priming (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b3_pilot_sizes_complete_cohort_and_removes_priming) ... ok\ntest_b4_entry_joins_effort_body_and_unknown_usage (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b4_entry_joins_effort_body_and_unknown_usage) ... ok\ntest_b5_net_difference_power_and_unrecovered_failure (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_b5_net_difference_power_and_unrecovered_failure) ... ok\ntest_canaries_have_reachable_shape_gates_and_valid_collision (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_canaries_have_reachable_shape_gates_and_valid_collision) ... ok\ntest_draft_contract (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_draft_contract) ... ok\ntest_keyed_confirmatory_env_and_no_ambient_openai_key (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_keyed_confirmatory_env_and_no_ambient_openai_key) ... ok\ntest_sensitivity_affinity_and_scope_are_explicit (tests.test_gpt6_lane_compression_ab_preregistration.CompressionPreregistrationTests.test_sensitivity_affinity_and_scope_are_explicit) ... ok\n\n----------------------------------------------------------------------\nRan 10 tests in 0.012s\n\nOK\n", + "evidence_class": "local structural/integrity validation" + }, + { + "phase": "blueprint_classification", + "command": "rtk env PYTHONDONTWRITEBYTECODE=1 TMPDIR=/var/tmp/claude-431 GIT_OPTIONAL_LOCKS=0 python3 -m unittest tests.test_blind_checkout.RepositoryClassificationTests -v", + "exit_code": 0, + "returned_output": "test_classified_texts_override_the_rules (tests.test_blind_checkout.RepositoryClassificationTests.test_classified_texts_override_the_rules) ... ok\ntest_every_blueprint_value_under_a_label_key_is_classified (tests.test_blind_checkout.RepositoryClassificationTests.test_every_blueprint_value_under_a_label_key_is_classified) ... ok\n\n----------------------------------------------------------------------\nRan 2 tests in 0.150s\n\nOK\n", + "evidence_class": "local structural/integrity validation" + }, + { + "phase": "publication_scan", + "command": "rtk env PYTHONDONTWRITEBYTECODE=1 TMPDIR=/var/tmp/claude-431 GIT_OPTIONAL_LOCKS=0 python3 scripts/validate.py --scan-file blueprints/gpt6-lane-compression-ab/PREREGISTRATION.md --scan-file blueprints/gpt6-lane-compression-ab/preregistration.json --scan-file tests/test_gpt6_lane_compression_ab_preregistration.py", + "exit_code": 0, + "returned_output": "{\"scanned_files\": 3, \"status\": \"passed\"}\n", + "evidence_class": "local structural/integrity validation" + }, + { + "phase": "post_record_whitespace", + "command": "rtk env GIT_OPTIONAL_LOCKS=0 git diff --check", + "exit_code": 0, + "returned_output": "", + "evidence_class": "local structural/integrity validation" + } + ], + "remaining_gates": [ + "runner-route", + "merged-profile", + "network-containment", + "header-capture", + "effort-detail", + "two-hop", + "effective-config", + "native-harnesses", + "native-controls", + "pilot-power-resources", + "usage-sensitivity" + ], + "registration_handoff": "Coordinator re-registers changed hashes in manifests/evidence.json; this repair never writes it.", + "failed_edit_attempts": [ + { + "phase": "document-render", + "exit_code": 1, + "returned_output_excerpt": "TypeError: replace() argument 2 must be str, not None", + "cause": "Historical source inventory contains a null URL; the temporary renderer did not exclude it. JSON was written, Markdown was not.", + "repair": "Filter source URLs to strings and regenerate both documents from read-only HEAD blob; no duplicate round records. Subsequent render exited 0 and 10 contract tests passed." + } + ], + "git_metadata_scope": "No mutations of the assigned repository git metadata and no agent git add/commit commands. Required upstream repository tests create their own disposable fixture repositories in TMPDIR.", + "sanitization": "Returned outputs preserve command exit codes and test summaries; worktree/home and random fixture paths are replaced with angle-bracket placeholders. The requested TMPDIR and read-only .git obstruction are recorded explicitly.", + "temporary_directory_observation": { + "requested": "/var/tmp/claude-431", + "dot_git_exists": true, + "dot_git_mode": "dr-xr-xr-x", + "workaround_performed": false, + "source": "tools/sota-convergence/blind_checkout.py:970-973 checks each ancestor for .git regardless of repository validity." + } + } +} diff --git a/manifests/evidence.json b/manifests/evidence.json index 736b1c3f3..730cbf681 100644 --- a/manifests/evidence.json +++ b/manifests/evidence.json @@ -6881,6 +6881,16 @@ "sha256": "7adef5c4578f4636a1596ff3e67f5dab43c051f6015c1d0e477f071579168462", "bytes": 16076 }, + { + "path": "blueprints/gpt6-lane-compression-ab/PREREGISTRATION.md", + "sha256": "a005e3193ba3b16592cfea273294a9210cc1f7ab5af327e2b6f24307ddb9f2fd", + "bytes": 52252 + }, + { + "path": "blueprints/gpt6-lane-compression-ab/preregistration.json", + "sha256": "34554a57569d8cbaff2fccf101b488eb743be19e14f6d7e623c90fc9ce582206", + "bytes": 272524 + }, { "path": "blueprints/memory-lifecycle-probe/README.md", "sha256": "c50ba6b05cba66095a357f1185429761b251462b9b91d1d23a51532b317bc042", diff --git a/tests/test_gpt6_lane_compression_ab_preregistration.py b/tests/test_gpt6_lane_compression_ab_preregistration.py new file mode 100644 index 000000000..4804d999b --- /dev/null +++ b/tests/test_gpt6_lane_compression_ab_preregistration.py @@ -0,0 +1,376 @@ +"""Local structural checks, not a model run or upstream harness acceptance. + +Reference seam: tests/test_token_e2e_preregistration.py at 0f76651d; +the #416 compaction draft supplies document structure only. +""" + +import hashlib +import json +from pathlib import Path +import unittest + + +ROOT = Path(__file__).resolve().parents[1] +BLUEPRINT = ROOT / "blueprints/gpt6-lane-compression-ab" +ENGINES = { + "session-dedup", "ccr", "lite", "rtk", "codex-responses", "headroom", + "relevance", "caveman", "aggressive", "llmlingua", "ultra", "omniglyph", +} +SAFE_LABELS = {"session-dedup", "ccr", "lite", "headroom"} +ROLES = {"builder", "reviewer", "researcher"} + + +def digest(value): + return hashlib.sha256(json.dumps( + value, sort_keys=True, ensure_ascii=True, separators=(",", ":"), + allow_nan=False, + ).encode("utf-8")).hexdigest() + + +class CompressionPreregistrationTests(unittest.TestCase): + def test_draft_contract(self): + # Fail first at the public document boundary, before either file exists. + self.assertTrue((BLUEPRINT / "preregistration.json").is_file(), + "preregistration.json does not exist yet") + self.assertTrue((BLUEPRINT / "PREREGISTRATION.md").is_file()) + data = json.loads((BLUEPRINT / "preregistration.json").read_text()) + prose = (BLUEPRINT / "PREREGISTRATION.md").read_text() + + with self.subTest(contract="honest draft"): + self.assertEqual(data["status"], "DRAFT") + self.assertFalse(data["frozen"]) + self.assertFalse(data["run_started"]) + self.assertFalse(data["execution_authorized"]) + self.assertIsNone(data["results"]) + self.assertIsNone(data["sealing"]["sealed_at"]) + self.assertIn("DRAFT", prose) + self.assertEqual(data["client"]["version"], "0.157.1") + self.assertEqual(data["client"]["profile"], "stack-worker") + self.assertEqual(data["client"]["effort"], "max") + + cells = {cell["id"]: cell for cell in data["cells"]} + with self.subTest(contract="clean control and crossed factors"): + self.assertEqual(set(cells), {"C", "D0", "D1", "A0", "A1"}) + self.assertEqual(cells["C"]["base_url"], "http://127.0.0.1:20128/v1") + self.assertEqual(cells["C"]["model"], "cx/gpt-6-astra-max") + self.assertEqual(cells["C"]["engines"], []) + for cell_id in ("D0", "D1", "A0", "A1"): + cell = cells[cell_id] + self.assertEqual(cell["classification"], "confirmatory") + self.assertEqual(cell["base_url"], "http://127.0.0.1:20129/v1") + self.assertEqual(cell["model"], "sharedgw/gpt-6-astra-max") + self.assertEqual(set(cell["engines"]), + SAFE_LABELS if cell_id[0] == "D" else ENGINES) + self.assertEqual(cell["compression_header"], + None if cell_id[0] == "D" else "allow-lossy") + self.assertEqual(cell["output_styles_on"], cell_id.endswith("1")) + self.assertFalse(data["factorization"]["header_off_is_clean_control"]) + self.assertEqual(set(data["exploratory"]["single_engine_order"]), + ENGINES - SAFE_LABELS) + self.assertEqual(data["factorization"]["keyed_live_zone"]["ttl_minutes"], 60) + self.assertEqual(data["factorization"]["keyed_live_zone"]["key_file_mode"], "0600") + self.assertFalse(data["exploratory"]["allows_promotion"]) + + with self.subTest(contract="frozen task packets and controls"): + self.assertEqual(set(data["task_sets"]), ROLES) + for role, packet in data["task_sets"].items(): + self.assertEqual(packet["sha256"], digest(packet["packet"]), role) + self.assertIn(packet["sha256"], prose) + tasks = packet["packet"]["tasks"] + self.assertGreaterEqual(len(tasks), 3) + self.assertEqual(len(tasks), len({t["id"] for t in tasks})) + for task in tasks: + self.assertTrue(task["prompt"]) + self.assertTrue(task["oracle"]) + self.assertTrue(task["source_ids"]) + if role != "builder": + examples = task["grader_controls"] + self.assertEqual(json.loads(examples[0]["output"]), + task["expected"]) + self.assertNotEqual(json.loads(examples[1]["output"]), + task["expected"]) + with self.assertRaises(json.JSONDecodeError): + json.loads(examples[2]["output"]) + controls = packet["packet"]["controls"] + self.assertEqual({c["kind"] for c in controls}, + {"known-pass", "known-fail", "malformed"}) + self.assertEqual([c["expected_accept"] for c in controls], + [True, False, False]) + self.assertNotEqual(controls[0]["input"], controls[1]["input"]) + self.assertNotEqual(controls[0]["input"], controls[2]["input"]) + canaries = data["canaries"] + self.assertEqual(canaries["sha256"], digest(canaries["packet"])) + self.assertIn("1234567890123456711", canaries["packet"]["numeric_tool_output"]) + self.assertIn("1.50", canaries["packet"]["numeric_tool_output"]) + self.assertGreater(len(canaries["packet"]["long_tool_output"]), 2000) + self.assertGreater(canaries["packet"]["long_tool_output"].index("TAIL_PIN="), 2000) + self.assertEqual(set(canaries["packet"]["roles"]), ROLES) + self.assertIn("multipart", canaries["packet"]["dedup_probe"]) + numeric = canaries["packet"]["numeric_engine_tool_output"] + self.assertGreater(len(numeric.encode()), 512) + lexical = json.loads(numeric, parse_int=str, parse_float=str) + self.assertEqual(lexical["integer"], "1234567890123456711") + self.assertEqual(lexical["decimal"], "1.50") + + with self.subTest(contract="measurement boundaries"): + metrics = {m["id"] for m in data["metrics"]} + self.assertTrue({"task_success", "apply_patch_failures", "verification_failures", + "schema_retries", "tool_output_recall", "tokens_in", + "tokens_cache_read", "tokens_reasoning", "compression_savings", + "latency", "cancellations"} <= metrics) + accounting = data["accounting"] + self.assertEqual(accounting["usage_authority"], { + "C": "20128.call_logs", "D0": "20129.call_logs", + "D1": "20129.call_logs", "A0": "20129.call_logs", + "A1": "20129.call_logs", + }) + self.assertEqual(accounting["savings_authority"], + "20129.call_logs.tokens_compressed") + self.assertFalse(accounting["add_savings_to_usage"]) + self.assertTrue(accounting["cache_read_is_input_subset"]) + self.assertTrue(accounting["reasoning_is_output_subset"]) + self.assertIsNone(accounting["output_column_verified_on_host"]) + self.assertIn("tokens_out", accounting["required_additional_columns"]) + self.assertIn("tokens_in - tokens_cache_read", accounting["uncached_input_formula"]) + self.assertEqual(accounting["total_billed_tokens_formula"], + "sum(tokens_in + tokens_out)") + self.assertTrue(accounting["include_failed_retried_cancelled_attempts"]) + + with self.subTest(contract="multi-turn transport and role gates"): + multi = data["multi_turn"] + self.assertGreaterEqual(multi["user_turns"], 3) + self.assertTrue(multi["tools_required_each_turn"]) + self.assertEqual(multi["aggregation_unit"], "task/session/user_turn") + self.assertTrue({"prompt_cache_key", "session_headers", "affinity_headers", + "store", "include", "encrypted_reasoning", "turn_2_tool_roundtrip"} + <= set(multi["required_checks"])) + self.assertEqual(data["role_policy"]["early_canary_roles"], ["builder"]) + self.assertTrue({"reviewer", "researcher", "judgment", "verification", "evidence"} + <= set(data["role_policy"]["control_until_pass"])) + self.assertTrue(data["role_policy"]["record_outputs_require_exact_preservation"]) + + with self.subTest(contract="statistics and stopping are specified"): + analysis = data["analysis"] + self.assertIsNone(analysis["planned_repetitions_per_role_cell"]) + self.assertIsNone(analysis["expected_repetitions_per_task_cell"]) + self.assertIn("independent", analysis["sampling"]) + self.assertIsNone(analysis["total_confirmatory_sessions"]) + self.assertEqual(analysis["success_ni_margin"], 0.05) + self.assertEqual(analysis["failure_ni_margin"], 0.05) + self.assertEqual(analysis["multiplicity"]["method"], "holm") + self.assertEqual(analysis["multiplicity"]["comparisons"], 12) + self.assertTrue(analysis["degenerate_interval_guard"]) + self.assertIn("bootstrap", analysis["upstream_functions"]) + self.assertIn("permutation_test", analysis["upstream_functions"]) + self.assertTrue(data["decision_rule"]["per_role"]) + self.assertEqual(data["decision_rule"]["objective"], + "minimum total billed tokens among non-inferior cells") + self.assertTrue(data["decision_rule"]["tie_rule"]) + self.assertEqual(data["decision_rule"]["minimum_absolute_combined_success_rate"], + 0.90) + budget = data["budget"] + self.assertIsNone(budget["total_token_cap"]) + self.assertIsNone(budget["allocations"]) + self.assertEqual(budget["max_concurrent_sessions"], 1) + self.assertTrue(budget["owned_process_cancellation"]) + self.assertTrue(budget["stop_conditions"]) + + with self.subTest(contract="source-backed hazards and unresolved acceptance"): + sources = {s["id"]: s for s in data["sources"]} + self.assertTrue({"harbor-codex", "promptfoo-responses", "scipy", "statsmodels"} + <= set(sources)) + self.assertEqual({h["id"] for h in data["hazards"]}, + {"H1", "H2", "H3", "H4", "H5", "H6"}) + for hazard in data["hazards"]: + self.assertTrue(hazard["checks"]) + self.assertTrue(hazard["verification_status"]) + self.assertIsNone(hazard["live_result"]) + for source_id in hazard["source_ids"]: + self.assertIn(source_id, sources) + for gate in data["open_host_gates"]: + self.assertIsNone(gate["passed"]) + self.assertTrue(data["anti_pattern_log"]) + # Every source-backed contract addition must resolve to the inventory. + def check_sources(value): + if isinstance(value, dict): + for key, item in value.items(): + if key in {"source_id", "source_ids"}: + for name in [item] if isinstance(item, str) else item: + self.assertIn(name, sources) + check_sources(item) + elif isinstance(value, list): + for item in value: + check_sources(item) + + check_sources(data) + + def read_contract(self): + return json.loads((BLUEPRINT / "preregistration.json").read_text()) + + def test_b1_per_hop_names_and_unqualified_runner_cannot_seal(self): + data = self.read_contract() + self.assertEqual(data["client"].get("model_hops"), { + "client_to_20129": "sharedgw/gpt-6-astra-max", + "20129_to_20128": "gpt-6-astra-max", + "control_C": "cx/gpt-6-astra-max", + }) + runner = data["harnesses"]["primary"] + self.assertEqual(runner["runner_choice"], "harbor-with-route-qualification") + self.assertEqual(runner["actual_model_argument"], "gpt-6-astra-max") + self.assertTrue(runner["sealing_blocked"]) + gate = runner["route_qualification"] + self.assertIsNone(gate["result"]) + self.assertIn("prompt-input", gate["checks"]) + self.assertIn("live-200-max", gate["checks"]) + self.assertIn("control-route-equivalence", gate["checks"]) + self.assertTrue(runner["overturn_condition"]) + self.assertEqual(len(runner["alternatives"]), 3) + + def test_b2_merge_launch_network_and_native_resume(self): + data = self.read_contract() + merge = data["client"].get("profile_merge", {}) + self.assertEqual(merge.get("inputs"), ["config.toml", "stack-worker.config.toml"]) + self.assertEqual(set(merge["forbidden_keys"]), {"profile", "profiles"}) + self.assertIsNone(merge["sha256"]) + self.assertEqual(set(merge["equivalence"]["compare"]), + {"resolved-config", "prompt-input"}) + self.assertIsNone(merge["equivalence"]["result"]) + argv = data["client"]["argv_contract"] + self.assertEqual(argv[:2], ["codex", "exec"]) + for flag in ("--dangerously-bypass-approvals-and-sandbox", + "--skip-git-repo-check", "--json", "--enable"): + self.assertIn(flag, argv) + self.assertNotIn("-p", argv) + self.assertIn("unified_exec", argv) + network = data["harnesses"]["primary"]["network"] + self.assertEqual(network["overlay"]["services"]["main"]["network_mode"], "host") + self.assertTrue(network["reachability_probe"]) + self.assertTrue(network["admin_api_exposure"]) + self.assertIsNone(network["result"]) + multi = data["multi_turn"] + self.assertEqual(multi["native_route"], "Harbor [[steps]]") + self.assertTrue(multi["resume_trajectory"]) + self.assertEqual(multi["task_turn_mapping"]["B-hello-multi-step-simple"], + ["create-file", "append-content", "exact-record"]) + + def test_b3_pilot_sizes_complete_cohort_and_removes_priming(self): + data = self.read_contract() + pilot = data.get("qualification_pilot", {}) + self.assertEqual(pilot.get("phase"), "pre-confirmatory-data") + self.assertTrue(pilot["excluded_from_confirmation"]) + self.assertEqual(set(pilot["cells"]), {"C", "D0", "D1", "A0", "A1"}) + self.assertTrue({"tokens", "wall_seconds", "paired_outcomes", "operational_events"} + <= set(pilot["measure"])) + self.assertIsNone(pilot["results"]) + self.assertTrue(data["budget"]["sizing_rule"]) + self.assertIsNone(data["budget"]["whole_run_wall_seconds"]) + self.assertEqual(data["analysis"]["priming_sessions"], 0) + self.assertEqual(data["analysis"]["cache_strata"], [1, 2, 3]) + + def test_b4_entry_joins_effort_body_and_unknown_usage(self): + accounting = self.read_contract()["accounting"] + self.assertIn("correlation_id", accounting["required_columns"]) + self.assertEqual(accounting["correlation_capture"]["header"], "X-Correlation-Id") + self.assertTrue(accounting["correlation_capture"]["source_ids"]) + self.assertIsNone(accounting["correlation_capture"]["qualified"]) + self.assertFalse(accounting["add_downstream_hop"]) + self.assertEqual(accounting["effort"]["field"], "reasoning.effort") + self.assertEqual(accounting["effort"]["hops"], [20129, 20128]) + self.assertFalse(accounting["effort"]["null_column_is_drift"]) + self.assertTrue(accounting["effort"]["request_detail_route"]) + missing = accounting["missing_usage"] + self.assertFalse(missing["zero_fill"]) + self.assertTrue(missing["request_size_is_not_total_bound"]) + self.assertTrue(missing["worst_case_ranking_required"]) + + def test_b5_net_difference_power_and_unrecovered_failure(self): + analysis = self.read_contract()["analysis"] + ni = analysis.get("paired_net_test", {}) + self.assertEqual(ni.get("method"), "scipy.stats.bootstrap") + self.assertTrue(ni["paired"]) + self.assertEqual(ni["alternative"], "less") + self.assertEqual(ni["statistic"], "mean(control_success - candidate_success)") + self.assertIn("beneficial", ni["exact_fallback"]) + self.assertEqual(analysis["multiplicity"]["unit"], "role-candidate intersection-union") + self.assertTrue(analysis["operational_failure"]["unrecovered_only"]) + power = analysis["power_simulation"] + self.assertGreaterEqual(power["target_power"], .80) + self.assertTrue(power["joint_paired_outcomes"]) + self.assertTrue(power["margin_null_calibration"]) + self.assertIsNone(power["selected_n"]) + self.assertIsNone(power["results"]) + + def test_keyed_confirmatory_env_and_no_ambient_openai_key(self): + data = self.read_contract() + key_name = "OMNIROUTE_FW_API_KEY" + for cell in data["cells"]: + if cell["id"] == "C": + continue + self.assertEqual(cell["principal"], "registered-lane-key") + self.assertEqual(cell["env_key"], key_name) + self.assertEqual(cell["live_zone_ttl_minutes"], 60) + auth = data["client"]["environment_contract"] + self.assertEqual(auth["extra_env"][key_name], "${" + key_name + "}") + self.assertIn("OPENAI_API_KEY", auth["must_be_unset"]) + self.assertFalse(auth["static_auth_headers"]) + + def test_canaries_have_reachable_shape_gates_and_valid_collision(self): + data = self.read_contract() + packet = data["canaries"]["packet"] + self.assertEqual({row["shape"] for row in packet.get("numeric_shapes", [])}, + {"shell", "mcp-text", "mcp-structuredContent", "custom-tool"}) + for row in packet["numeric_shapes"]: + self.assertIsNone(row["native_result"]) + self.assertTrue(row["qualification"]) + messages = packet["dedup_messages"] + self.assertEqual(len(messages[0]["content"]), 2) + self.assertIn("EARLIER_USER_PIN=", messages[0]["content"][1]["text"]) + self.assertEqual(messages[1]["content"], messages[2]["content"]) + self.assertNotIn(messages[1]["content"], messages[0]["content"][1]["text"]) + self.assertEqual(packet["dedup_assert_unchanged"], "messages[0].content[1].text") + self.assertTrue(data["oracle_contract"]["score_exactness_when_engine_not_applied"]) + self.assertTrue(data["oracle_contract"]["engine_coverage_separate"]) + self.assertEqual(data["factorization"]["sharedgw_wire_api"], "responses") + + def test_sensitivity_affinity_and_scope_are_explicit(self): + data = self.read_contract() + ranges = data["accounting"].get("sensitivity", {}) + self.assertEqual(ranges.get("cache_ratio"), [0.0, 1.0]) + self.assertEqual(ranges["output_ratio"], [1.0, 10.0]) + self.assertFalse(ranges["actual_price_claim"]) + affinity = data["multi_turn"]["affinity"] + self.assertEqual(affinity["connection_pin"], "x-omniroute-connection") + self.assertNotIn(affinity["connection_pin"], affinity["session_headers"]) + self.assertTrue(affinity["account_spread_by_arm"]) + self.assertTrue(affinity["body_derived_live_zone_fallback"]) + self.assertEqual(data["decision_rule"]["scope_ceiling"], "named synthetic task domains only") + self.assertEqual(data["decision_rule"]["review_research_style_scope"], "JSON-only records") + + def test_all_review_findings_have_cited_round_dispositions(self): + data = self.read_contract() + round_record = data.get("repair_round_20260927", {}) + rows = round_record.get("findings", []) + self.assertEqual({r["id"] for r in rows}, + {f"B{i}" for i in range(1, 6)} | + {f"M{i}" for i in range(1, 11)} | + {f"m{i}" for i in range(1, 9)}) + self.assertEqual(len(rows), 23) + for row in rows: + self.assertIn(row["outcome"], {"fixed", "fixed by gating", "declined"}) + self.assertTrue(row["reason"]) + self.assertTrue(row["citations"]) + self.assertTrue(row["verification"]) + self.assertFalse(round_record["model_inference_performed"]) + self.assertFalse(round_record["git_metadata_written"]) + self.assertEqual(round_record["installed_hashes_checked"], 17) + self.assertEqual(set(round_record["remaining_gates"]), + {gate["id"] for gate in data["open_host_gates"]}) + prose = (BLUEPRINT / "PREREGISTRATION.md").read_text() + for row in rows: + self.assertIn(f"| {row['id']} | {row['outcome']} |", prose) + for gate in round_record["remaining_gates"]: + self.assertIn(f"**{gate}**", prose) + + +if __name__ == "__main__": + unittest.main()