Skip to content

New WSL local models: settle the generation and embedding slots by their preregistered measurement - #647

Merged
seathatflowsinourveins merged 11 commits into
mainfrom
foundation/local-model-results-20261003
Oct 3, 2026
Merged

seathatflowsinourveins merged 11 commits into
mainfrom
foundation/local-model-results-20261003

Conversation

@seathatflowsinourveins

Copy link
Copy Markdown
Owner

Scope

  • What this PR changes, in one or two sentences: it settles the new WSL's two local-model slots by their preregistered measurement of 2026-10-02/03 (local-generation-model: Swift-1.5-Qwen3.8-27B IQ3_S on Ollama 0.35.0 as swift-iq3s-s2o-64k; embedding-model: qwen3-embedding:0.6b as qwen3-embedding-8k). It publishes the measurement's sanitized records and two receipts, records both settlements in the definitive manifest (with the manifest, the definitive-defaults tables and the handbook regenerated), makes the two install-plan rows installable from pinned sources, and adds the decision record docs/decisions/2026-10-03-new-wsl-local-models.md.
  • Base commit: d2777ee76a6ca14e020856fcd76421ceebc1e4f0
  • Lane: lane:foundation. The manifests/evidence.json change is the hot-file protocol's re-registration, which does not by itself make a PR shared (docs/lanes.md, "Labels and PRs").
  • Owned paths touched:
    • evidence/artifacts/new-wsl-local-models-20261002/ (new): the preregistration's published copy, its hash chain and copy notes, the two receipts, the per-case and per-query records, the frozen scripts, the Modelfile reconstruction and files.json.
    • evidence/artifacts/new-wsl-definitive-defaults-20261001/: settlements.json, assemble_manifest.py, controls.py and render_tables.py; definitive-manifest.json is generated.
    • evidence/artifacts/new-wsl-install-plan-20261002/: install-plan.json, owners.json, install.sh, accept.sh, check_plan.py, models/*.Modelfile, README, SOURCES and VALIDATION.
    • evidence/artifacts/new-wsl-handbook-20261001/receipt.json, tests/test_new_wsl_definitive_defaults.py (class LocalModelAcceptance), and .gitignore (two exceptions to *.jsonl).
    • docs/decisions/2026-10-03-new-wsl-local-models.md (new). Generated: the tables of docs/decisions/2026-10-01-new-wsl-definitive-defaults.md, docs/new-wsl-handbook.json and docs/new-wsl-handbook.md.
    • Shared hot file, in the last commit only: manifests/evidence.json.

SOTA sources

  • Model server, Ollama v0.35.0: https://github.com/ollama/ollama, lightweight tag v0.35.0 on cc4069396f3ad2c370c53eed2e4a42ac13adab84. Each command of the two rows is read at that commit (evidence/artifacts/new-wsl-install-plan-20261002/SOURCES.md, "The two local-model rows"):
    • docs/cli.mdx#L85 and server/images.go#L1082: pull and the layer digest check.
    • cmd/cmd.go#L2422 and docs/modelfile.mdx#L120, #L126, #L146: create from a GGUF file, a relative FROM, and num_ctx.
    • envconfig/config.go#L112, #L277 and manifest/paths.go#L29: the model store and the OLLAMA_NUM_PARALLEL default.
    • docs/api.md#L48, #L1351 and #L1659: generate with think false, list, and /api/embed.
    • server/routes.go#L928-L1151, #L981-L984 and #L3226-L3227: the embed handler answers 404 for a missing model and pulls nothing.
  • Generation model: https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF at revision d74895bbe5db4bec1e0024e7cc87d59c02d7631a. The file is Swift-1.5-Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf: LFS sha256 1333c6ea70ef348d4ac6d62732772e8ad6571ac5b3754c14ed54f1a0d904a786, 11,771,546,912 bytes. It is served with the Modelfile lines of the Ollama library model qwen3.8:27b, which evidence/artifacts/new-wsl-local-models-20261002/reconstruct_s2o_modelfile.py rebuilds independently.
  • Embedding model: the Ollama library model qwen3-embedding:0.6b (Q8_0), manifest digest ac6da0dfba84a81fdbfbaf330198c33cd77c4cdfc53e8bc50eb581914a15621d. Upstream: https://github.com/QwenLM/Qwen3-Embedding. Model card: https://huggingface.co/Qwen/Qwen3-Embedding-0.6B/blob/97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3/README.md#L36.
  • Upstream evaluation harnesses (none self-written):
  • Deciding measurement and preregistration:
    • The refuting Claude critic's text: evidence/artifacts/new-wsl-final-architecture-20261002/critics/added-critics-result.json (sha256 ed444353fae463da52f5f461fb5975eb5d607260d6c219407e77567255b69622).
    • evidence/artifacts/new-wsl-local-models-20261002/PREREGISTRATION-local-models.md: published copy sha256 454f879d03d61fa5a8e9f0d877e62cd94878abcd32f8f09b16ac9514492c3a0b; a 16-entry hash chain whose final text is d161feaacd6b3082b3adf37e80a368b58bd6f68205efd399102b8d84d0d0f78b.
    • The other model family's read: Client configuration for the new distribution: wire only what the definitive manifest installs (map, tool, F9) #608 (comment)
  • Evidence registration: docs/lanes.md, "Hot-file protocol". manifests/evidence.json is main's copy, with this branch's files re-registered by scripts/host_receipts.py register_file and its two receipts re-added. During the rebase, main's copy was kept by a merge driver given on the command line, using the custom-driver mechanism in https://git-scm.com/docs/gitattributes ("Defining a custom merge driver").

Evidence-class table

Claim Evidence class Command / receipt
Generation slot: S2o (Swift IQ3_S on Ollama 0.35.0, swift-iq3s-s2o-64k) made 200 of 200 first-pass valid tool calls on 100 BFCL simple_python cases through the chat and responses surfaces, against 175 of 200 for the control gpt-oss:20b. The difference is 0.125, with a 95% paired bootstrap interval of 0.065 to 0.19. Both arms ran entirely on the GPU with the embedder resident local_integration (native model runs on one workstation in a throwaway distribution, one run per case and arm; not the destination) evidence/artifacts/new-wsl-local-models-20261002/part-a-receipt.json (sha256 55ac6ff05f4437dffb2623117c71fb8d6f786d71e63a44b46973b145b1d29577, kind native_model_e2e)
Embedding slot: qwen3-embedding:0.6b (qwen3-embedding-8k) scored nDCG@10 0.49827 against 0.41394 for EmbeddingGemma on CQADupstackUnixRetrieval (difference 0.084331, interval 0.067738 to 0.100656). Nemotron-3-Embed-1B scored 0.5014, but its interval of -0.012194 to 0.018141 includes zero and it needs a second model server, so it does not take the slot local_integration (same scope) evidence/artifacts/new-wsl-local-models-20261002/part-b-receipt.json (sha256 6f5dfbf83c5116e38b42d38ddd5fdbb1a27cf23b5ef7e8ac12d4071c9173a541, kind native_model_e2e)
Both decision statistics re-derive from the published per-case and per-query records local_integration (re-derivation from the records on 2026-10-03; nothing was re-measured) The four commands in evidence/artifacts/new-wsl-local-models-20261002/README.md; outputs in part-a/bootstrap.txt and part-b/bootstrap.txt
The definitive manifest settles both rows: their state becomes measurement, the decision count installed goes from 56 to 58 and definitive stays at 31. The record's tables and the handbook are current local_integration assemble_manifest.py --check, render_tables.py --check and build_new_wsl_handbook.py --check, each exit 0; controls.py exit 0 (41 mutations killed, 53 tests OK)
The two rows' install and acceptance commands follow Ollama v0.35.0's documented behaviour: the pull digest check, create from a relative GGUF FROM, a per-model num_ctx, and a 404 from /api/embed for a missing model, with no pull source_review evidence/artifacts/new-wsl-install-plan-20261002/SOURCES.md, "The two local-model rows", read at cc4069396f3ad2c370c53eed2e4a42ac13adab84
The four acceptance programs and the local-model-server row's after_sign_in check pass on stand-ins and fail on each planted defect synthetic (stub ollama and curl programs, a scratch HOME and model store) tests/test_new_wsl_definitive_defaults.py, class LocalModelAcceptance, within the 142-test run (exit 0)
The plan's inventory, owners, scripts and manifest agree: 69 rows, of which 40 installed, 3 measurement-only and 26 not installed local_integration (static) check_plan.py exit 0; bash -n on install.sh and on accept.sh, exit 0 each
After its FROM line, the plan's Swift Modelfile equals the measured one local_integration (recorded by the branch on 2026-10-03; not rerun for this PR) reconstruct_s2o_modelfile.py --check returned 0 (SOURCES.md)
manifests/evidence.json is main's copy plus this branch's 63 file entries (44 new, 19 re-hashed) and 2 receipts local_integration scripts/validate.py exit 0: 9464 hashed files, 189 receipts
The 65 changed files add no private-content match local_integration scan_private.py, the coordination lane's scanner (outside the repository): its 13 matches in three files are the same, file by file, on origin/main and at the pre-rebase head

Not established (decision record, "What is not established"):

  • No command of either row has run as a plan row anywhere.
  • Nothing is accepted on the destination distribution.
  • GPU ownership between the workstation's and the destination's model servers is left to the lifecycle design.
  • The reranking condition was not run.
  • One host and one task per slot.

Local commands run

From the branch's worktree at head b9d0bf629883e7719c76e0655aa59c4adcbd3e92, 2026-10-03 (checks from 08:28:15Z):

$ git fetch origin && git -c core.attributesFile=<scratch>/attrs -c merge.keepmain.driver=true rebase --onto origin/main fd111e59a748
  # <scratch>/attrs holds the single line: manifests/evidence.json merge=keepmain
exit 0: 9 commits replayed onto d2777ee76; the 3 evidence-only commits became empty and were dropped
$ python3 -c '...; host_receipts.register_file(Path("."), p)' <the branch's 63 registered paths>   # docs/lanes.md hot-file protocol
exit 0
$ cd evidence/artifacts/new-wsl-definitive-defaults-20261001 && python3 -B assemble_manifest.py --check
exit 0: definitive-manifest.json is current
$ cd evidence/artifacts/new-wsl-definitive-defaults-20261001 && python3 -B controls.py ../../..
exit 0: 41 mutations killed; Ran 53 tests, OK; all files restored
$ cd evidence/artifacts/new-wsl-definitive-defaults-20261001 && python3 -B render_tables.py --check ../../../docs/decisions/2026-10-01-new-wsl-definitive-defaults.md
exit 0: the record's tables are current
$ cd evidence/artifacts/new-wsl-install-plan-20261002 && python3 -B check_plan.py
exit 0: OK: 69 rows: 40 installed (38 by the default run, 2 only when named), 3 measurement-only, 26 not installed; 74 commands and 61 acceptance entries agree with the scripts
$ bash -n evidence/artifacts/new-wsl-install-plan-20261002/install.sh; bash -n evidence/artifacts/new-wsl-install-plan-20261002/accept.sh
exit 0; exit 0
$ python3 -B scripts/build_new_wsl_handbook.py --check
exit 0: {"status": "passed"} (docs/new-wsl-handbook.json 974e78af..., docs/new-wsl-handbook.md 3e9ad2be...)
$ python3 -B -m unittest tests.test_new_wsl_definitive_defaults tests.test_new_wsl_handbook tests.test_new_wsl_profile
exit 0: Ran 142 tests, OK
$ python3 -B scripts/validate.py
exit 0: {"components": 69, "hashed_files": 9464, "profiles": 4, "receipts": 189, "status": "passed"}
$ python3 -B scripts/validate_convergence.py --all-recorded
exit 0: record[0] to record[25]: valid
$ python3 -B scripts/component_matrix.py --check; python3 -B scripts/new_host_grand_list.py --check
exit 0: {"rows": 32, "status": "checked"}; exit 0: {"status": "passed", "layers": 32, "winners": 66}
$ scan_private.py <checkout> <the 65 changed files>
exit 1: 13 matches (handbook JSON and MD: 1 home path and 5 e-mail each; evidence.json: 1 home path), the same file by file on origin/main and at the pre-rebase head: 0 new

Decision record

docs/decisions/2026-10-03-new-wsl-local-models.md. It names the deciding measurement, the alternatives (gpt-oss:20b, the Bonsai and PrismML arms, EmbeddingGemma, Nemotron-3-Embed-1B, WeMM-Embedding-2B) and the comparisons that would overturn each slot.

Host evidence

Not applicable: nothing under evidence/hosts/ changes.

Checklist

  • New/changed GitHub Actions are pinned to a full commit SHA with a version comment (no floating tags). (none changed)
  • New/changed workflows declare top-level permissions: contents: read (or a narrower, explicitly justified addition). (none changed)
  • No secrets are printed, logged or committed; no new required secret was added without a documented owner.
  • No new paid hosting, subscription or billing surface was introduced.
  • Peer-owned untracked files and worktrees were preserved (not deleted, moved or overwritten).

Merge slot: after session 0c's #630, #637, #626 and #628, in the order agreed on #608.

🤖 Generated with Claude Code

@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Oct 3, 2026
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-03T08:47:37.854648Z b9d0bf6 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b9d0bf6298

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread evidence/artifacts/new-wsl-install-plan-20261002/accept.sh
Comment thread evidence/artifacts/new-wsl-install-plan-20261002/install.sh
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

SOURCE ACCEPT for bounded historical evidence publication at b9d0bf6, against published main d2777ee. The consequential Astra/max source read returned that disposition; the native review process closed 0. This permits publishing the limited measurements. Final architecture, destination and SDK role qualification remain open.

The decision expressly labels both results as local integration checks on one workstation/throwaway distribution, one run per case and arm, and excludes destination acceptance or a claim to the best available model. The settlements become measurement; neither becomes definitive. The displayed installed count is explained as a decision count, and the new plan rows have not run as plan rows.

The publication map has 35 public copies (32 declared identical, three declared sanitizations) and 196 private metadata entries. I independently checked all 35 published pinned Git blobs against their declared byte counts and SHA256: 35 match, 35 native exits 0. Map SHA256 4e77359503133255d0dc30f44e6c577839bfa6124d1146f6ebb0ddff423f1893. This does not reread or qualify the private originals. The previous 175-artifact map is a distinct contract; the README discloses the later 55,649-file copy-out's unchecked checksum list and private native-output locators.

Independent processing of the exact head/main registry originals confirms 9420→9464 file rows, 187→189 receipts, 44 added rows, 19 changed bindings and no removed rows; all previous receipts and all 26 convergence records remain. Head registry SHA256 cc980de682ae4421d42f93ad2f3635e4ad7300094906b4f46fc09b748cd64594. These are consistency checks, not execution acceptance. The review's final registry command failed because my prompt supplied the wrong cached filename; I corrected the locator and checked the retained main original (SHA256 60959d71327bde82547afbab87b689532b71b1514fccb524e9c8f9493e179197). The original failed read is retained. I do not infer an unregistered-decision blocker from a policy that does not impose that requirement.

Failed generation arms, missing failure text, co-residency limits and unrun reranking remain evidence limits. native_model_e2e is interpreted with the receipts' Local integration check class, under the acceptance policy. Receipt publication does not create a new native run or upstream test. The bounded Astra read did not cover the full Part B receipt or all cancellation/cross-layer assertions; its acceptance does not independently authenticate those runtime claims.

Pinned Ollama Modelfile documentation supports the relative-GGUF/context configuration mechanism. The pinned keep-alive implementation supports the documented residency difference. Actual loading, served dimensions, GPU ownership/restoration, residency, destination install and broader role quality still require their own native evidence. No evaluated model, benchmark, GPU or destination operation was run for this review. The source review itself was one requested Astra/max model run; actual served model/effort, billing and upstream request count remain unknown. Required CI still governs merge eligibility.

seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
…d local-model rows

ManifestRuleTests counted 38 installing rows (36 of the 64-row plan plus the two consensus rows). Settling
local-generation-model and embedding-model adds two installing rows, so the plan installs 40 (check_plan.py: 40
installed, 38 by the default run, 2 only when named). The set equalities of the same test already held; only the
count was stale. Found by the hosted validate run of #647 (run 37110319677).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

SOURCE ACCEPT carried to ae90cb2, one commit directly atop root-accepted b9d0bf629883e7719c76e0655aa59c4adcbd3e92. The previous bounded historical-evidence scope and all destination/runtime/role limits remain.

The immutable delta changes only tests/test_new_wsl_client_config.py:438: comment plus expected plan-row count 38→40, accounting for the two settled local-model rows. The preceding set-equality assertions are unchanged. Diff +3/−2; no behavior, evidence or model result changes.

Both heads bind the registry to Git blob 293ae3af4ac1e900e292a802d63dca3036ac6a12, 2,523,755 bytes / SHA256 cc980de682ae4421d42f93ad2f3635e4ad7300094906b4f46fc09b748cd64594. The modified test is not a hash-registered payload. Every other accepted artifact/map/source input is unchanged, so the matching 35-copy checks and bounded source judgment are reused rather than rerun.

The old-head validate failure, completed09:08:08, remains a failure. At09:28:59 the new-head validation was in progress and macOS jobs queued. The owner's 146-tests/validator0 report is distinct from a completed hosted check. Seven original native API captures returned0; the immutable compare and exact changed lines were read independently. No tests, models or host operations were rerun for this small source repair.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

FINDINGS / required macOS gate open, exact #647 head ae90cb2b30756a6a7d93025618a62f503009710f, base d2777ee76a6ca14e020856fcd76421ceebc1e4f0, final native guard11:28:11Z. Source-publication acceptance does not establish a passing native CI gate.

The existing mac job111173581589/run37112688327 failed11:20:17Z; step20's full suite failed11:20:06Z. Its original artifact11273275413, bound to this run and exact head, contains 9849 tests /1737.068s /two failures /1328 skips. The only failures are generation-files L1133 and embedding-files L1149, each the synthetic expected-state case returning1 when0 is required. These are integration fixtures, not actual model runs.

Both execute the exact install-plan acceptance programs, embedding L846 and generation L1569: sha256sum --check --status, followed by a manifest grep. The new fixture's L1075 checks executable presence only. The pinned maintained capability probe L793–804 explicitly handles macOS /sbin/sha256sum rejecting those GNU options and uses a scratch native shasum fallback L905–906.

This is a source-supported portability lead. The failed fixture captures then discards inner stdout/stderr atL1105–1106, so original CI does not identify whether checksum or later grep failed. Do not claim the precise inner cause has been reproduced.

Concrete owner repair: reuse that existing capability probe/fallback for the synthetic fixture and expose failed subprocess diagnostics. Preserve its positive and planted-negative oracles and the production WSL acceptance programs. Then obtain native macOS evidence for the repaired head. No GPU/provider/model operation is needed for these synthetic checks. Root hands off the owner paths and has changed no Claude-lane files.

One authorized isolated Linux native run at this exact head passed all five LocalModelAcceptance tests, exit0,0.189s, including both positive cases and their negatives. It used an owned clean worktree and restricted environment; this confirms Linux fixture validity and is not a macOS reproduction. No broad suite was rerun by root.

Root independently verified21 non-auth original command stdout/stderr bindings and seven pinned source payload identities, actual artifact run/head and ZIP-member equality. The packet retains22 commands:20exit0, two retrieval exit1s (gh run view --log saw incomplete run metadata; direct job-log API needed its supported escape-sequence flag); recovered direct retrieval exit0. Artifact ZIP291299B/SHA256 bfa099366fdd01c5e07e522dd901ad0fbbd6b011c0900c1fd7543f6fa5945f72; actual full log1842862B/SHA 137b1d0c1147c1acf523f8eb1de5e9a504bb0b2b8d1d24b5303681424c07c3d7. Capture manifest25955B/SHA e24176eead553f2e2eacb1c87e2b1504045806c6d75584dce93fb9ef7f74062e. The original failure and retrieval errors remain retained; Linux success does not replace them. Source identity and original logs remain separate from installed-model and destination qualification.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Codex exact-head SOURCE ACCEPT for the checksum-shim and diagnostic repair; FINDINGS before landing at febc6e87383cd9a4dfea6e4dce6e3d68349adf71, against held ae90cb2b30756a6a7d93025618a62f503009710f, captured main 4ced2923063db6a6dcafa9f25af5ee05a4153c75.

The repaired fixture uses the existing sha256sum_checks_like_gnu capability predicate and, when needed, the scratch shasum -a 256 shim, matching the maintained bootstrap precedent. The existing Bash subprocess and all five case definitions are preserved; assertion diagnostics now include exit, stdout and stderr. The delta changes no production acceptance program or model evidence. I did not rerun these controls in the shared dirty checkout.

Landing finding: the registry still carries eight foreign bindings that differ from current main: .github/workflows/catalog-freshness.yml, AGENTS.md, docs/github-automation.md, docs/harness-defaults.md, tools/sota-convergence/README.md, tools/sota-convergence/build_manifest.py, tools/sota-convergence/extract_layers.py and tools/sota-convergence/github_freshness.py. Refresh from main and re-register the owned files under the shared-file protocol before merge. This is a stale evidence-binding finding, not a Git-file deletion claim. All 9,420 main paths remain present; all 187 main receipt entries and 26 convergence records are preserved, with two owned receipts added.

The original Mac run's two positive fixture failures remain retained. Their original inner stderr was not supplied, and source inspection does not authenticate the proposed checksum cause or establish that the repaired Mac run passes. The earlier passing Linux controls retain their synthetic/integration scope. A successful native repaired Mac result is still needed for that claim.

Custody: CODEX-PR647-REPAIR-20261003T120402Z/manifest.json, 16,572 bytes, SHA256 ee4506fe9ef144411d7a53db402cf30a569e3f786d18e6e41c489e8d11c2269e; all 17 native captures exited 0, stdout/stderr bindings and seven pinned Git blobs checked. Repaired test Git blob b0941052293fc59d5bed34e586361cb2652630a8, SHA256 06dedeebec7f01da17db62ce2392e8a569d598da3d49dddf889a95f5326eb30d. Initial oversized-contents responses had encoding:none and failed decoding; that failure is retained, and immutable Git-blobs API recovery supplied the verified .decoded.source registry payloads. No model, GPU, provider, credential or client-configuration operation was performed.

seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
…d local-model rows

ManifestRuleTests counted 38 installing rows (36 of the 64-row plan plus the two consensus rows). Settling
local-generation-model and embedding-model adds two installing rows, so the plan installs 40 (check_plan.py: 40
installed, 38 by the default run, 2 only when named). The set equalities of the same test already held; only the
count was stale. Found by the hosted validate run of #647 (run 37110319677).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the foundation/local-model-results-20261003 branch from febc6e8 to bd52c9d Compare October 3, 2026 12:52
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Root exact-head SOURCE ACCEPT for the landing-carry repair.

Head bd52c9d, base/captured main 4ced292. Fresh native guard13:01:14 confirms OPEN/non-draft. This clears the eight foreign-binding landing finding in the febc source read.

Root compared complete, non-truncated Git trees at the accepted febc6e8 and current head. Their literal delta is12 paths: the registry and11 carried current-main paths. Each of those11 matches main's Git identity exactly. All65 PR file paths other than the registry retain their accepted old Git identities. The three-dot API comparison uses the older d277 merge base and is retained as that comparison, rather than treated as a literal old-to-new diff.

All9401 foreign current-main file bindings now match;9464 registry membership and current-main order are preserved. The eight previously stale file rows are the only old-to-new registry-row changes. Owned rows,189 receipt rows (including all187 current-main rows) and26 convergence records remain unchanged. The checksum test is byte-identical to accepted febc: Git blob b0941052293fc59d5bed34e586361cb2652630a8,78877B/SHA256 06dedeebec7f01da17db62ce2392e8a569d598da3d49dddf889a95f5326eb30d.

The prior checksum capability/shasum fallback source acceptance carries forward. Old macOS failure output and synthetic Linux fixture results remain distinct from this head's native macOS/host/GPU qualification. Required CI must pass at this head before owner merge; this read did not rerun tests or prove the cause of a historical failure.

Original source custody: complete held/head trees SHA256 b20dac7d9e6154703780394253dc73fa73d037ff314dc982ec1e8e7bd53e6e61 / 36a2a0d437f33bd4ea8093aa735aa32095bb3f46142795f90ea984b49ce7aa7e; decoded registry2523755B/SHA256 90abbf49237696530c2d560ba6e8b5eb03dfd243778681cb9f590af6bf06ff59, Git f5ee0dad000c0424313cc182a652ca32c9cb6531. Public immutable originals and native exits retained in root's CODEX-ROOT-PR647-LANDING-20261003T1300Z packet. No Claude-owned paths or private client/auth files were edited or read.

Scout and others added 9 commits October 3, 2026 12:18
…records and receipts

A sanitized copy of the preregistration (one replacement: the throwaway
distribution's name) with its hash chain and copy notes; receipts for the
generation slot (Part A and its extension) and the embedding slot (Part B);
the per-case validity files of the four scored A2 runs and the per-query
nDCG@10 files of E1 to E3, from which both decision statistics reproduce
unchanged with the published scripts; the frozen scorers, runners and steps;
a reconstruction of the measured Swift Modelfile from the Ollama library's
blobs (sha256 8911245e... with the measurement's FROM line); files.json lists
every file of the private measurement folder by sha256.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…l by their measurement

Two settlements in the form of the model-server settlement, each with the
preregistration copy and its receipt. Both slots are rows that the
convergence decisions add, so the assembler now settles such a row right
after those decisions (it must be a split of the rounds; its resolution
stays as recorded; its installed job stays owned by one row). The record's
tables show a returned measurement's settlement also on a split row. Tests
follow the new state (58 rows install, 6 measurement, 5 split); one control
refuses a settlement aimed at an added row that is not split. Manifest,
tables, handbook and the handbook receipt's output hashes regenerated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… sources

local-generation-model downloads the Swift IQ3_S file at its Hugging Face
revision with its sha256, places the repository's two Modelfiles beside it and
creates swift-iq3s-s2o and swift-iq3s-s2o-64k (context 64,000 as the model's
own parameter). embedding-model pulls qwen3-embedding:0.6b, stops unless the
server lists the pinned library manifest digest, and creates
qwen3-embedding-8k (context 8,192). No server-wide context is set. Both rows
create their models through the running model server, which the plan does not
start, so they install and are checked only with --only.

Acceptance: post_install reads files only (the placed Modelfiles, the library
manifest's digest, the created models' layers); service_health shows each
model's own context and makes one short generation or one embedding call.
check_plan.py now checks the dispatch of rows installed only when named in
both directions; LocalModelAcceptance runs the four programs against stand-ins.
README, SOURCES and VALIDATION say what was read and checked, and that no
command of the two rows has run as a plan row anywhere. GPU ownership between
the workstation's model services and the destination's server is left to the
lifecycle design, as a stated limitation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…preregistered measurement

docs/decisions/2026-10-03-new-wsl-local-models.md records the deciding
measurements as the manifest names them, the preregistration and its
amendments, the arms, gates and rules of Part A and Part B, both results with
their decision statistics, the four deviations and the recorded defects, the
other model family's read (pull request 608, comment 5964331790) and the
artifact map it asked for, what the install plan does, what is not established
(the destination's acceptance, physical co-residency in pb2, one host and one
task each, and the wider measurement the critic's own text named), the sources
by sha256 and the overturn conditions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…placed, server smoke check on the settled embedder)

- .gitignore: exceptions to *.jsonl for the four per-case validity files and the three per-query
  nDCG@10 files, which both decision statistics reproduce from; the seven files are now committed
  (hashes unchanged, as files.json and both receipts record them).
- scripts/steps/W-partb-window.sh: the published copy named the throwaway distribution on line 18;
  it now writes '<the throwaway distribution>' (quoted, still parses with bash -n). files.json marks
  it changed with both hashes; COPY-NOTES gives 32 identical and 3 changed copies, corrects the
  byte-identical list and records the distribution-name scan the first publication lacked.
- files.json: the 231 files are described as listed; the copy-out added to the private folder
  afterwards is listed by its checksum list, with the measured Swift Modelfile (sha256 8911245e...).
  README and the decision record no longer say that file was not retained.
- part-a receipt and the decision record: the control's A1b footprint is 12.97 GB
  (12,968,494,366 bytes, 12.08 GiB), not GiB.
- Receipts' files.json hash, settlements, manifest, handbook and its receipt regenerated; the
  manifest's state counts are unchanged.
- Install plan: the local-model-server row's after_sign_in smoke check is one /api/embed call to
  qwen3-embedding-8k instead of `ollama run embeddinggemma`, which would pull the losing arm; a
  stand-in test covers it. README and SOURCES: the embedder's 32,768-token load, the private step
  scripts listed by hash, and the measured server's NUM_PARALLEL/KEEP_ALIVE against v0.35.0's defaults.
- Decision record: cites the Codex lane's binding read of the artifact map (file bindings only;
  an exact-head verdict is still required).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…smoke check

The repair round cited server/routes.go:975-977 at cc4069396f for "/api/embed answers 404 for a missing
model". Those lines answer 404 only when getExistingName fails, which happens when listing the manifests
fails; for a model the store does not have it returns the name unchanged, GetModel fails with the
manifest's not-found error, and handleScheduleError answers 404 "model %q not found, try pulling it first"
(routes.go:981-984 and 3226-3227). EmbedHandler (928-1151) calls no pull. accept.sh's comment, the
local-model-server row's notes and SOURCES.md now cite those lines; the stand-in test's error answer is
that message. The check itself is unchanged.

check_plan.py exit 0 (69 rows; 74 commands and 61 acceptance entries agree); bash -n accept.sh exit 0;
LocalModelAcceptance: 5 tests OK.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d local-model rows

ManifestRuleTests counted 38 installing rows (36 of the 64-row plan plus the two consensus rows). Settling
local-generation-model and embedding-model adds two installing rows, so the plan installs 40 (check_plan.py: 40
installed, 38 by the default run, 2 only when named). The set equalities of the same test already held; only the
count was stale. Found by the hosted validate run of #647 (run 37110319677).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… not GNU's, and report a failed program's output

The required macOS job of run 37112688327 (job 111173581589) failed the expected-state
cases of test_the_generation_files_check and test_the_embedding_files_check with exit 1.
The fixture discarded the program's output, so the inner cause was not recorded. Both
programs run `sha256sum --check --status`, and setUpClass only checked that a sha256sum
was on PATH.

The fixture now reuses tests/test_adoption_bootstrap.py's probe sha256sum_checks_like_gnu
and, where it is false, writes the same scratch sha256sum that runs `shasum -a 256` as that
module's run_install_pin; setUpClass requires shasum only then. run_program returns the
finished process, and each case's assertion message carries the program's exit code,
stdout and stderr. The case tables, the expected exits and the install plan's programs
are unchanged. The test file's re-registration in manifests/evidence.json is in the
branch's last commit (docs/lanes.md, hot-file protocol).

Linux only: the module passes, and with a stub sha256sum that rejects --check --status
first on PATH (a simulation, not macOS) all five LocalModelAcceptance tests pass through
the shasum shim, where the previous head failed the same two expected-state cases.
Native macOS evidence comes from CI on this head.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… skip the server smoke check until the embedder exists

install.sh's model_server_answers, which both local-model rows run before their
first command, accepted any server that answered `ollama ls`. It now also stops
the row unless GET /api/version reports 0.35.0, the measured server
(docs/api.md:1824 and server/routes.go:2023 at cc4069396f), before anything is
downloaded, pulled or created.

The local-model-server row's after_sign_in check calls qwen3-embedding-8k, which
only the embedding-model row creates, and step F9 runs that stage without any
model row, so it failed there every time. accept.sh now prints skipped, which is
not a pass, until that model's manifest is in the store the embedding row's
post_install check reads; the check itself is unchanged.

No command or acceptance entry of install-plan.json changes; its notes, the
README, SOURCES, VALIDATION and the decision record describe both repairs.
LocalModelAcceptance runs accept.sh's stage itself (skipped, 0, 1), and
ModelServerGuard runs the guard against stub ollama and curl (0.35.0 passes;
another version, no version, a failed request and no server fail; both rows run
it before their first command).

The local-model evidence README notes that M6-a1b.py's docstring repeats A1's
order while its code and record follow amendment 2's; the hash-bound copy is
unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the foundation/local-model-results-20261003 branch from bd52c9d to e471dee Compare October 3, 2026 16:21
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Exact-head FINDINGS on e471dee7036ad13abe62d36fc5251633140f385d, reread against accepted bd52c9d8540f6d31a71aa7d7e06480c92b48ad4c and main/base ecea28654a835fff2cc3651bab77ca0e46b9bec5. This is source review; the live plan rows, destination GPU qualification and this head's CI/landing remain separate gates.

Three of the five review concerns have source corrections:

  • The canonical F9 smoke check now prints skipped before requesting embeddings when the selected derived-model manifest is absent. That skip is explicitly not acceptance. accept.sh:403.
  • Both model rows call the server guard before their first mutation; it requires the running server to report 0.35.0. install.sh:136, supported by Ollama's pinned version route. Acceptance-time server replacement remains an explicit unverified condition in VALIDATION.md:255.
  • The A1b README now discloses the actual arm → embedder → long-prompt order and explains the contradictory docstring while keeping the historical script unchanged. README.md:41. Its private-record assertion was not independently reread here.

Two historical evidence concerns remain open. Please preserve the scripts that ran and add an impact/qualification disclosure bound to their existing original evidence; editing an as-run script would not repair a past trial.

  1. Part B precondition exit is masked. inside pipes the remote command into clean without pipefail; precheck returns PIPESTATUS[0] of that wrapper, which is the cleaning pipeline's status rather than the remote precheck's exit. A remote precondition failure with retained output can therefore be reported as success. E471 keeps these lines unchanged, and the captured thread has no owner impact response. State which independent, contemporaneous observation establishes the required precondition for each admitted Part B run, or mark the affected qualification unknown/failed. W-partb-window.sh:18, original review thread.
  2. G0 checksum output does not gate the trials. sha256sum -c feeds grep -vc and sed, and the following warm-up/scored loop proceeds without checking a checksum success condition. The reported changed-file count is diagnostic output, not a fail-closed prepared-file guard. E471 preserves this source and the captured thread has no owner impact response. Bind any claim that the prepared files were unchanged to the original checksum result and run inputs, and retain uncertainty where that evidence is unavailable. M11-g0-run.sh:12, original review thread.

Custody and landing: 90 old→new delta paths consist of 80 exact main carries and 10 owned modifications; the other 56 paths of the complete 66-path PR retain their accepted Git identities. All 35 publication-map members, both historical receipts and all three reviewed historical scripts are unchanged. I verified 28 complete repository originals, two pinned upstream blobs and 46 original native stream pairs, all exits 0, against byte/SHA256/Git bindings. No measurement, model, private pilot or test was executed for this read.

All 9,478 foreign registry row payloads and 191 existing receipts in order are preserved, as are the 29 convergence records and unowned metadata. Two adjacent foreign decision rows are reordered into sorted order: the SDK-worker and PR320-retirement decisions, the same explicit normalization proposed in #666. This is a narrow, identified order change, not row loss. The owner should carry that accepted normalization consistently with #666 and check whole-result sortedness in addition to payload preservation against actual main before landing. Source acceptance does not substitute for the required current-head checks, GitHub thread disposition or landed-tree observation.

Scout and others added 2 commits October 3, 2026 13:10
…heir impact

W-partb-window.sh's precheck (lines 18 and 28) returned the status of the
cleaning pipeline rather than the remote precheck's, so a failed precondition
could not stop a Part B window. The new section binds each admitted window's
precondition to the precheck's own printed lines in raw/_gpu-window.txt (sha256
42c3758a...): pb1 at 23:13:16Z and 23:13:27Z, pb2 at 23:48:48Z and 23:48:55Z,
each recording resident=[] llama-server pids=[]. M11p-precheck.sh exits 0
exactly when those two printed values are empty. Window pb1 ran the earlier
revision amendment 4a hashed (688722f7...), which was not retained, so the
defect is stated for the retained revision only.

M11-g0-run.sh line 12 prints its checksum count without testing it. The section
binds "prepared files unchanged" to each G0 record's printed count of 0
(raw/M11-g0-S1.txt 4849c444..., raw/M11-g0-S2.txt 09046e46...,
raw/M11-g0-S1b.txt e85547ce...) and notes that every arm that ran G0 failed it,
so no decision rests on a G0 pass.

Both scripts stay as they ran. Answers PR #647 review threads r4172365849 and
r4172365852.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…o receipts on main's copy

Start from main's manifests/evidence.json (ecea286) and apply this branch's
entries through the docs/lanes.md hot-file protocol: host_receipts.register_file
for the 63 files the branch registers (44 new, 19 re-hashed, now including the
review repairs to install.sh, accept.sh, install-plan.json, the install plan's
README, SOURCES and VALIDATION, the local-model evidence README and
tests/test_new_wsl_definitive_defaults.py) and the two receipts[] entries
re-added at the end (part-a-receipt.json and part-b-receipt.json of
new-wsl-local-models-20261002). convergence_records[] is unchanged. This commit
replaces the previous evidence-only commits bd52c9d and e471dee. Against
e471dee it changes one entry: the local-model evidence README, re-hashed
after the commit before this one added its known-defects section (sha256
deeac61a..., 13,179 bytes). files[] keeps e471dee's order.

Main's copy at ecea286 lists two files out of path order
(docs/decisions/2026-10-03-omniroute-sdk-worker-0160.md after
docs/decisions/2026-10-03-retire-pr320-loki-denominator-host-receipts.md), which
scripts/validate.py rejects; scripts/evidence_manifest.py --write, the command
validate.py names, swapped those two entries and changed nothing else.
component_matrix.py --write and new_host_grand_list.py --write leave the
manifest unchanged; scripts/validate.py passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the foundation/local-model-results-20261003 branch from e471dee to 4c033f1 Compare October 3, 2026 17:12
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Codex root exact-head SOURCE ACCEPT WITH NOTES for the two historical-script impact disclosures at 4c033f1eeccb256a40834d1f876fdf81dfd6fef1, from reviewed e471dee7036ad13abe62d36fc5251633140f385d, against observed main/base ecea28654a835fff2cc3651bab77ca0e46b9bec5. This closes the two disclosure findings in my E471 read.

The added README section states both defects and their consequences: the window wrapper masked the remote precondition exit; G0 printed a checksum diagnostic without gating the subsequent runs. It preserves the as-run scripts and names the printed observations on which the owner bases the impact assessment. It also expressly records that pb1's wrapper revision was not retained, so that revision's defect cannot be established from source. This is the needed historical disclosure, not a retroactive repair of an enforced gate.

Evidence scope: I verified the public source, artifact metadata and bindings. I have not read the four private witness bodies or independently authenticated their printed observations. The statement at L68 that the defect leaves no qualification unknown is therefore the owner's assessment from those named records, not new native qualification by this root. The same boundary applies to the reported three G0 failures and unchanged-file diagnostic. No new run or causal/performance/default acceptance follows.

The complete E471→4c03 delta changes only README and its registry row: README13,179B, SHA256deeac61a846f9d82ea4caf1994de37dcca074a92466c129139a705c1f7ebffc2, Git blob5edf1f492df2779a772a6d1f0a79e1920fb020e5. Root independently matched both full changed Git/SHA originals. All9,541 registry rows keep their order and full fields except that one binding; all193 receipts and29 convergence records are identical. The other64 PR paths and as-run artifacts are unchanged. The inherited adjacent foreign ordering normalization remains the separate #666 landing concern, not a new payload change here.

Original native captures remain retained, including the initial GraphQL1 from querying #608 as an issue and its corrected PullRequest recovery0. The fresh root guard returned0 and this exact head/main. No root test, model, GPU, service, private-log or runtime execution occurred; CI, prospective merge preservation and installed/default qualification remain separate.

@seathatflowsinourveins
seathatflowsinourveins merged commit 54a96ff into main Oct 3, 2026
25 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the foundation/local-model-results-20261003 branch October 3, 2026 18:12
seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
…utation runs in manifests/evidence.json

Hot-file protocol (docs/lanes.md): this is the branch's last commit, on main d2fc380 (#647, #661), and takes main's
manifests/evidence.json. It re-registers the closure record, the receipt and docs/harness-defaults.md, and the 71
files under evidence/artifacts/frozen-variant-guard-mutations-20261003/. component_matrix --write and
new_host_grand_list --write changed nothing else.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
…k is regenerated on top of main's #647 version (only this PR's research-state.json input-hash line) and the receipt's two output hashes are refreshed (standing ACK); registry carries main's rows plus the owned bindings

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
…otocol)

#647 settled local-generation-model and embedding-model by their preregistered
measurement; this branch adds the wave-2 records, interim installs, statusline
and the client-configuration map. Both survive. 15 paths conflicted.

Regenerated with the repository's generators, not hand-merged:
- definitive-manifest.json: evidence/artifacts/new-wsl-definitive-defaults-20261001/assemble_manifest.py
  (90 slots, 31 definitive, installed 59, interim 3; by_state resolved 23,
  measurement 6, split 5; consensus 6), then --check.
- docs/decisions/2026-10-01-new-wsl-definitive-defaults.md: its only conflict was
  inside the generated tables; render_tables.py --write, then --check.
- docs/new-wsl-handbook.json and .md: scripts/build_new_wsl_handbook.py --write.
- evidence/artifacts/new-wsl-handbook-20261001/receipt.json: no generator writes
  it; the two output digests come from the builder's own digest function (as
  --write printed them); generator, profile and inventory were unchanged.

Hand-merged, keeping both sides:
- install.sh: both header notes; both functions (interim_acknowledged,
  model_server_answers); --list takes this branch's code-search line and main's
  embedding-model line; the dispatch loops take this branch's selected lists and
  main's named-only list (local-generation-model, embedding-model).
- accept.sh: both header notes; the skipped-slot loop drops all five rows that
  now install (code-search, memory-owner, context-supply, embedding-model,
  local-generation-model), in plan order.
- install-plan.json, owners.json: revised_at_utc set to the merge's time
  (2026-10-03T18:26:54Z); the plan's status keeps main's sentence, then this
  branch's. Rows auto-merged (disjoint: 13 rows here, 3 there).
- README.md: counts from the merged plan (70 rows: 44 installed, 42 by the
  default run with three interims and two only when named, 3 measurement-only,
  23 not installed; post_install 28/15/1, service 10, after sign-in 7; route
  counts with model-server 2); the default-run rule keeps the interim clause
  and main's named-only sentence.
- README.md, SOURCES.md, VALIDATION.md: both new sections, in the order written
  (the local-model sections first written 05:25Z, wave 2 at 10:16Z), which
  VALIDATION's opening line requires. VALIDATION's opening line now says 70
  rows and names the Wave 2 section, and Wave 2's checks list records the
  merged plan's check_plan.py line.
- tests/test_new_wsl_client_config.py: 44 planned rows.
- tests/test_new_wsl_definitive_defaults.py: the merged counts above; both
  sides' test classes (StatuslineInstallAndAcceptance; LocalModelAcceptance,
  ModelServerGuard).

manifests/evidence.json: 54a96ff's copy (git checkout --theirs), then
host_receipts.register_file for the branch's 55 registered files (17 of them
re-hashed by this merge); receipts[] and convergence_records[] are main's.
component_matrix.py --write and new_host_grand_list.py --write change nothing;
scripts/validate.py passes (9542 hashed files).

check_plan.py: OK: 70 rows: 44 installed (42 by the default run, 2 only when
named), 3 measurement-only, 23 not installed; 94 commands and 66 acceptance
entries agree with the scripts. bash -n install.sh and accept.sh: exit 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
…-pin-20261003

Refresh onto main cac8700 (#651, #628, #647, #661, #648) under the docs/lanes.md
hot-file protocol.

- Handbook: main's docs/new-wsl-handbook.{md,json} and evidence/artifacts/new-wsl-handbook-20261001/receipt.json
  (#647) are the base; scripts/build_new_wsl_handbook.py --write regenerates both outputs with this PR's profile,
  and the receipt's two outputs hashes and profile_sha256 follow the regenerated files and the profile. Nothing
  else in the receipt changes.
- manifests/evidence.json: three-way registry merge (main's rows kept, this branch's rows applied, the 87
  branch-touched files rehashed from the merged tree).
- #628's Codex SDK worker pins openai-codex and openai-codex-cli-bin 0.160.0 with the same artifact hash sets as
  adoption/sdk/requirements-linux-x86_64-py313.lock and cites rust-v0.160.0 a956835d; it changes none of
  manifests/stack.json, adoption/pins-linux-x86_64.json, tools/adoption/apply_codex_lane.py or adoption/templates.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant