Skip to content

New WSL clean-install selection per foundation layer (blind judges on upstream evidence, adversarial critics) - #589

Merged
seathatflowsinourveins merged 4 commits into
mainfrom
foundation/new-wsl-clean-install-selection-20261001
Oct 1, 2026
Merged

seathatflowsinourveins merged 4 commits into
mainfrom
foundation/new-wsl-clean-install-selection-20261001

Conversation

@seathatflowsinourveins

Copy link
Copy Markdown
Owner

Scope

  • What this PR changes: records the new WSL's clean-install selection for the 20 foundation layers and the base distribution, chosen blind on upstream repository evidence (workflow wf_c2377ebb-65c), as a decision record and an evidence folder. No catalog, profile, pin or tool changes.
  • Base commit: 798ac445307e2cd8eba6e74d7722ac0e16da02c7
  • Lane: lane:shared (foundation content; manifests/evidence.json is a shared hot file, last commit only)
  • Owned paths touched: docs/decisions/2026-10-01-new-wsl-clean-install-selection.md, evidence/artifacts/new-wsl-clean-install-selection-20261001/ (29 files with the record), manifests/evidence.json.

Result: 15 layers selected, 5 install their arms for a head-to-head on the new WSL (isolation's container slot, semantic RAG, durable memory, token efficiency, the container engine), base distribution Ubuntu 26.04.1 LTS with 24.04.5 LTS as the fallback. The record lists what the critics changed and the 7 checked facts that did not hold (all in reasons, none overturning a pick). The 12 us-equities layers run the same method on the GPT lane after 2026-10-03T17:14Z, owned by the trading lane.

SOTA sources

  • Method: blind packets and symmetric evidence as designed in docs/decisions/2026-10-01-u11-merit-neutral-selection.md (accepted for implementation at revision 4, 89424e36, by the cross-family and the Claude reviewer); position, verbosity and self-preference bias of model judges per Zheng et al. 2023 (arXiv:2306.05685v4) and OpenAI's "Evaluation best practices" (https://developers.openai.com/api/docs/guides/evaluation-best-practices).
  • Each pick's evidence: the upstream URLs the judges and critics read, listed per pick in selection.json (install_source) and in the record.
  • The user's merit rule: decision 5 of docs/decisions/2026-10-01-definitive-sota-wsl-program.md.

Evidence-class table

Claim Evidence class Command / receipt
The per-layer picks and their reasons source_review by model judges with adversarial critics selection.json; the full returns are kept privately with their sha256 in the folder README
Criteria and prompts were frozen before the run structural_validation preregistration.json (recorded 2026-10-01T17:35:24Z; run started 17:35:59Z)
Any pick works on the new WSL none: not run the comparisons and stage 2 on the new distribution

Local commands run

$ python3 scripts/validate.py                                       # passed: 69 components, 8938 hashed files, 184 receipts
$ python3 scripts/evidence_manifest.py --check                      # passed
$ python3 -m unittest -q tests.test_validate tests.test_ecosystem_manifest   # Ran 138 tests, OK
$ git diff --check origin/main HEAD                                 # clean

Files rows 8,909 to 8,938 (29 own), none lost; receipts 184 and convergence records 26 unchanged.

Decision record

docs/decisions/2026-10-01-new-wsl-clean-install-selection.md.

Host evidence

Not relevant.

Checklist

  • No workflow or action changes.
  • No secrets are printed, logged or committed (the full returns stay out of the repository because the scanner reads their 40-hex commit ids as tokens).
  • No new paid hosting, subscription or billing surface was introduced.
  • Peer-owned untracked files and worktrees were preserved.

🤖 Generated with Claude Code

… upstream evidence, adversarial critics

Workflow wf_c2377ebb-65c: 11 Claude Opus judges read blind packets (requirement, candidates as name and repository,
seeded shuffle; the project's own receipts excluded) with criteria frozen and hashed before launch, and 3 adversarial
critics re-checked facts and stronger candidates. 15 foundation layers are selected, 5 install their arms for a
head-to-head on the new WSL, and the base distribution is Ubuntu 26.04.1 with 24.04.5 as the fallback. The record
lists what the critics changed and the 7 facts that did not hold. Source review, not a measured result.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins seathatflowsinourveins added the lane:shared Touches files owned by both lanes; needs both lanes' acknowledgement label Oct 1, 2026
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Trading lane: manifest part acknowledged at 3fd48722 (base 798ac445), read in code. manifests/evidence.json goes from 8,909 to 8,938 files, with 29 added and none re-pinned or dropped. No added row is a trading path. receipts (184) and convergence_records (26) are unchanged, and none of the 30 changed files is under blueprints/us-equities, catalogs/us-equities or catalogs/landscape/us-equities. The decision record's scope note on the 12 us-equities layers matches the plan: I run the same method unchanged on the GPT lane after the pool resets on 2026-10-03.

…rison arm listed, field coverage stated

The user asked for each layer to use its best stack with no overlap. ownership.json gives each tool exactly one
owning layer; other layers list it under uses. Ten overlaps are resolved: the two clients and Worktrunk get one owner,
one container-engine comparison and one local-model-server comparison replace two each, Phoenix is the trace sink only
(evaluations belong to Inspect AI), systemd comes with the distribution, mise installs uv, and four tool pairs keep
distinct roles. From the Gate A owner's review: the compare layers list every arm their deciding comparison names,
deja-vu has its upstream install command, the record states which candidates the packets did not hold, that all 14
agents were one model family, that the Harbor token-tool run was left out of its packet, and that the profile pins
and verifies every quick-start install command.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the foundation/new-wsl-clean-install-selection-20261001 branch from 3fd4872 to 4576aa5 Compare October 1, 2026 19:15
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Trading lane: re-acknowledged at 4576aa5c, read in code. manifests/evidence.json goes from 8,909 to 8,939 files: 30 added (the new one is ownership.json), none re-pinned or dropped, and no trading path. receipts (184) and convergence_records (26) are unchanged.

Scout and others added 2 commits October 1, 2026 15:19
… cross-family status; one owner per tool at the edges

From the Codex lane's independent audit: the picks are source-review recommendations, not merit winners; each layer's
comparison on the new distribution selects (program decision 5). The judges' close-call flags (19 of 21 layers) are
carried into the table, the seven failed facts get their supported replacements with sources, Letta's exclusion moves
to its supported criterion, and the record states that one model family judged and how the GPT-6.1 Sol reviews of
PR #575 relate. From the Gate A owner's check: eleven quick-start installs, isolation's arms no longer repeat the
engines that hosting-services owns, the code-RAG and engine comparisons run before the server and boundary
comparisons, and the trading rule cites the trading lane owner's acceptance. The boundary with the us-equities layers
is assigned (DuckDB, DVC, MLflow, pandera, agent-retrieval-bench). One anti-pattern row records the mistake.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ifests/evidence.json (hot-file protocol, last commit)

Files 8,909 -> 8,939 (+30 own: the packets, criteria, prompts, preregistration, packet builder, selection manifest,
ownership map and README, and the decision record); docs/harness-defaults.md re-registered if listed. Receipts and
convergence records unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the foundation/new-wsl-clean-install-selection-20261001 branch from 4576aa5 to 2c178e5 Compare October 1, 2026 19:20
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Trading lane: re-acknowledged at 2c178e54, read in code.

Manifest. 8,909 to 8,939 files: 30 added, 1 re-pinned, none dropped, and no trading path. receipts and convergence_records are unchanged.

The two passages of ownership.json, read as the trading owner.

  • boundary: correct. DuckDB with Parquet goes to storage-compute, DVC to identity-provenance, MLflow to evaluation-experiments and pandera to data-quality-orchestration. agent-retrieval-bench stays with the foundation's quality-evaluation as a candidate task set.
  • trading_rule: correct in substance, with one wording correction for the next edition update (no need to move this head). "its earlier trading selections that are now foundation-owned (Inspect AI, gitleaks, Grype, the observability stack) become uses of the foundation layers" can be read as the trading rows using gitleaks and Grype from the foundation. The foundation selected neither. Please make it: "...become uses of the foundation layers' selections: today Inspect AI, the observability stack, betterleaks and trufflehog for secrets, and the foundation's CI scanning; gitleaks and Grype leave the trading rows."

@seathatflowsinourveins
seathatflowsinourveins merged commit 20ea4ae into main Oct 1, 2026
25 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the foundation/new-wsl-clean-install-selection-20261001 branch October 1, 2026 19:58
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Cross-family half: run and recorded (2026-10-01, Claude session)

The blind GPT-6.1 Sol run this record requested has run on the same frozen packets, criteria and prompts, all sha256-verified against preregistration.json.

Result:

  • Agreement with Claude's picks: exact on 2 layers (native-clients, agent-sdks), overlap on 19, differ on 0.
  • GPT critics: upheld 4, revised 1 (workers), undetermined 16. They ran 156 fact checks. The 1 that failed refuted a judge's claim that Canonical's listing lacked the 26.04.1 WSL image: the critic found that image, dated 2026-08-27, with signed checksums.
  • Final catalog: the frozen rule is applied as written, so every pick both families made stands as its layer's pick for the new WSL. All 21 judged layers have standing picks: two_family_pick 2, shared_pick 3, partial_comparison 16. Memory has no clean winner yet: ai-memory is the blind pick both families made, but the only measurement on record (source host, descriptive, without the reranker) favors agentmemory (recall_all@5 0.821 against 0.496), Hindsight is unmeasured, and the Memory layer S3 r7: symmetric merit rule, maintenance gate, current pins, token ledger, RAG head-to-head (draft, not frozen) #526 head-to-head decides.

For the trading lane's run after 2026-10-03T17:14Z:

  • cross-family/run_xfam.py and build_record.py take the packets unchanged.
  • Usage this time, as returned by Codex: 20.6M input tokens (18.5M cached) and 360K output across the 14 processes, about 1.3M input per judge.
  • The Codex CLI pool answered at 20:41Z. The limit message above is the cloud code-review quota.
  • The ownership.json wording correction you asked for belongs to the next edition update and is listed on Grand catalog finalization (2026-09-23): live board #140.

seathatflowsinourveins added a commit that referenced this pull request Oct 3, 2026
… the rule extensions

Repairs the cross-family review findings on #595 (GPT-6.1 Sol, read of
025c489).

- Finding 2: the final catalog is now the record of the blind GPT-6.1 Sol
  half of the clean-install selection and its comparison with the blind
  Claude record (#589). It states at the top that it is not an install
  list and names the definitive manifest (#602) as the install record,
  without reading it (no --check coupling). The standing-picks and
  challengers framing, the gate ledger, install commands and the judges'
  deciding-comparison texts are gone; rows carry neutral pick sets
  (named by both halves, by one half only) and evidence classes.
- Finding 1: the record and the generated rule section say the fold
  extends the frozen agreement rule in two places (equal sets with unequal
  statuses; packet-name matching of distro names), and every judged row
  carries the class the rule's text gives as written next to the
  generator's.
- Finding 4: mentioned() matches only a full owner/name; regression test
  through fold() with non-empty arms, and --check runs against real stale
  files in a temporary root.
- Finding 5: agreement-rule.txt's sha256 is recorded in the JSON and
  --check fails when it changes without regeneration; tested.
- Finding 3: the record and the cross-family README disclose the timing
  and inventory limits (no per-attempt timestamps, the start time resting
  on the private run log, the instruction file's hash only from the
  next-day probe).

selection-gpt.json, agreement-rule.txt, the preregistration files and the
judges' outputs are unchanged. Hot-file protocol: this branch's files are
re-registered and the component matrix, the new-host grand list and the
final catalog regenerated with their --write commands.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
… through #622

Each new gate row cites an in-repository evidence file that its PR added and
quotes that file's own facts:

- wsl-retrieval-retirement-20261003 (#622)
- new-wsl-definitive-defaults-20261001 (#589, #591, #602)
- new-wsl-local-model-server-20261002 (#598)
- new-wsl-distro-recipe-20261002 (#593)
- new-wsl-install-plan-20261002 (#606, #607)
- new-wsl-client-configuration-20261002 (#608)
- mac-memory-qualification-closure-20261002 (#603)
- sdk-useful-task-preparation-20261002 (#609, #612)
- two-host-architecture-20261002 (#610)

The three lanes, the six workers and the 48 existing gates are unchanged;
recorded_at_utc comes from date -u and meaning is rewritten for this
checkpoint. The state.json row of manifests/evidence.json is re-registered with
host_receipts.register_file (docs/lanes.md hot-file protocol). The generation
holds 117 entities with the workflow adapter unconfigured, 127 with ten Dagu
runs (cap 128).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
… through #622

Each new gate row cites an in-repository evidence file that its PR added and
quotes that file's own facts:

- wsl-retrieval-retirement-20261003 (#622)
- new-wsl-definitive-defaults-20261001 (#589, #591, #602)
- new-wsl-local-model-server-20261002 (#598)
- new-wsl-distro-recipe-20261002 (#593)
- new-wsl-install-plan-20261002 (#606, #607)
- new-wsl-client-configuration-20261002 (#608)
- mac-memory-qualification-closure-20261002 (#603)
- sdk-useful-task-preparation-20261002 (#609, #612)
- two-host-architecture-20261002 (#610)

The three lanes, the six workers and the 48 existing gates are unchanged;
recorded_at_utc comes from date -u and meaning is rewritten for this
checkpoint. The state.json row of manifests/evidence.json is re-registered with
host_receipts.register_file (docs/lanes.md hot-file protocol). The generation
holds 117 entities with the workflow adapter unconfigured, 127 with ten Dagu
runs (cap 128).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
… through #622

Each new gate row cites an in-repository evidence file that its PR added and
quotes that file's own facts:

- wsl-retrieval-retirement-20261003 (#622)
- new-wsl-definitive-defaults-20261001 (#589, #591, #602)
- new-wsl-local-model-server-20261002 (#598)
- new-wsl-distro-recipe-20261002 (#593)
- new-wsl-install-plan-20261002 (#606, #607)
- new-wsl-client-configuration-20261002 (#608)
- mac-memory-qualification-closure-20261002 (#603)
- sdk-useful-task-preparation-20261002 (#609, #612)
- two-host-architecture-20261002 (#610)

The three lanes, the six workers and the 48 existing gates are unchanged;
recorded_at_utc comes from date -u and meaning is rewritten for this
checkpoint. The state.json row of manifests/evidence.json is re-registered with
host_receipts.register_file (docs/lanes.md hot-file protocol). The generation
holds 117 entities with the workflow adapter unconfigured, 127 with ten Dagu
runs (cap 128).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Oct 3, 2026
19 units and 43 open foundation slots of the new-WSL definitive defaults, each asked as a neutral function question,
with every candidate of each unit's field (177 contenders: the frozen #589 packets plus the runtime rows' fields).
Every candidate gets a dossier from its upstream source at its release tag (README read in full, code, packaging,
tests and CI, README claims checked against code), verified once and repaired once on a failed check. Both families
then decide per unit in two orders, a critic per family checks both deciders, and contested slots go to adjudicators
of both families in both A/B orders; a slot adjudication does not settle keeps two finalists and the named
measurement. Every model call runs in a clean room: claude -p in safe and restricted mode, and codex exec with an
isolated CODEX_HOME on the OmniRoute gateway; clean-room.json records the probes. preregistration.json holds the
sha256 of every frozen file, recorded at 2026-10-02T03:04:22Z before the first dossier.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Oct 3, 2026
… through #622 (#640)

Each new gate row cites an in-repository evidence file that its PR added and
quotes that file's own facts:

- wsl-retrieval-retirement-20261003 (#622)
- new-wsl-definitive-defaults-20261001 (#589, #591, #602)
- new-wsl-local-model-server-20261002 (#598)
- new-wsl-distro-recipe-20261002 (#593)
- new-wsl-install-plan-20261002 (#606, #607)
- new-wsl-client-configuration-20261002 (#608)
- mac-memory-qualification-closure-20261002 (#603)
- sdk-useful-task-preparation-20261002 (#609, #612)
- two-host-architecture-20261002 (#610)

The three lanes, the six workers and the 48 existing gates are unchanged;
recorded_at_utc comes from date -u and meaning is rewritten for this
checkpoint. The state.json row of manifests/evidence.json is re-registered with
host_receipts.register_file (docs/lanes.md hot-file protocol). The generation
holds 117 entities with the workflow adapter unconfigured, 127 with ten Dagu
runs (cap 128).

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:shared Touches files owned by both lanes; needs both lanes' acknowledgement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant