Skip to content

Token-efficiency profile accepted with the three code-navigation tools; task-to-model routing record (unit A4) - #540

Merged
seathatflowsinourveins merged 41 commits into
mainfrom
claude/sota-defaults-a4-20260930
Oct 1, 2026
Merged

seathatflowsinourveins merged 41 commits into
mainfrom
claude/sota-defaults-a4-20260930

Conversation

@seathatflowsinourveins

@seathatflowsinourveins seathatflowsinourveins commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Scope

  • What this PR changes: accepts the token-efficiency adoption profile as the selected practice (2026-09-30) with its 14 pinned components, keeps jcodemunch-mcp, codebase-memory-mcp and ast-grep as the SubagentStart carrier's on-demand lanes (not profile members until each has a reviewed pin on both platforms plus a host receipt: the bootstraps fail closed on an unpinned component), and adds the task-to-model routing record (one table: task class, client, model, effort, enforcement point today; no automatic router; OmniRoute D04/D06/D11 stay held until unit D3's promptfoo A/B; gpt-6.1-sol named only as pending the D4 qualification). Restates the README profile row (accepted; cells describe the 14 pinned components) and the regenerated grand list.
  • Base commit: 11227bfd (origin/main at rebase)
  • Lane: lane:foundation
  • Owned paths touched: adoption/manifest.json (token-efficiency profile label and recipe_paths), adoption/README.md (profile row), docs/decisions/2026-09-30-task-model-routing.md (new), docs/token-practice.md (pointer), tests/test_adoption_status.py, tests/test_task_model_routing.py (new); generated: catalogs/landscape/new-host-grand-list.json, docs/new-host-grand-list.md; manifests/evidence.json (re-registration only, last commit).
  • Frozen surfaces (Gate A): none touched; the profile's component set is unchanged, so tools/token-e2e/freeze_snapshot.py captures no new item.
  • Round 2 (after the macOS gate failed on the round-1 head): the profile is back to its 14 pinned components; the new test test_every_profile_component_has_a_pin_on_both_platforms enforces it. Token stack winner: one full stack chosen on recorded evidence, provisional until Gate A (decision record) #508 ("Token stack winner", head b7fcc219) agrees for jCodeMunch (owner only once wired) and codebase-memory-mcp (not a member); ast-grep is out only for want of a pin. The record's overturn condition names the follow-up: a reviewed pin in both pin files plus a host receipt per platform, then they join the profile.
  • Acceptance covers the selection and its routing only; no host is accepted (each host still runs the coverage check, the pinned-version check and one native call per tool per client).

SOTA sources

Evidence-class table

Claim Evidence class Command / receipt
The profile resolves with 13 commands and 5 recipes on this host local_integration python3 scripts/adoption_status.py --profile token-efficiency --json: exit 0, prerequisites_present
Profile membership (14 pinned), landscape selection, README row, routing record binding local_integration python3 -m unittest tests.test_token_stack_cards tests.test_adoption_status tests.test_adoption_docs_consistency tests.test_task_model_routing: 172 tests OK (1 data-dependent skip); tests.test_adoption_bootstrap_macos tests.test_adoption_bootstrap tests.test_adoption_launchd: 311 OK (38 environment skips)
Generated reports current local_integration python3 scripts/new_host_grand_list.py --check passed; python3 scripts/build_ecosystem.py --check passed
The record's quotes and the profile-to-record binding local_integration tests/test_task_model_routing.py: 8 OK; failing-first against the old manifest
macOS bootstrap refusal, #508 rows, freeze-snapshot behaviour source_review adoption/bootstrap-macos.sh:30-31; #508 head lines 85/86/103; tools/token-e2e/freeze_snapshot.py:2240-2300
Profile acceptance structural validation artifact consistency, not execution or adoption (docs/acceptance-evidence-policy.md)

No unchanged upstream tests are part of this change. No provider run or model call was made.

Local commands run

$ HOME=<scratch> TMPDIR=/dev/shm/ntci-tmp python3 -m unittest tests.test_token_stack_cards tests.test_adoption_status tests.test_adoption_docs_consistency tests.test_task_model_routing
Ran 172 tests   OK (skipped=1)
$ python3 scripts/new_host_grand_list.py --check
{"status": "passed", "layers": 32, "winners": 66}
$ python3 scripts/build_ecosystem.py --check
{"status": "passed", ...}
$ python3 scripts/validate.py
{"components": 69, "hashed_files": 8515, "profiles": 4, "receipts": 176, "status": "passed"}

Decision record

docs/decisions/2026-09-30-task-model-routing.md

Host evidence

No files under evidence/hosts/ changed.

Checklist

  • No GitHub Actions changed.
  • No workflows changed.
  • No secrets printed, logged or committed.
  • No paid hosting, subscription or billing surface.
  • Peer-owned untracked files and worktrees preserved.

🤖 Generated with Claude Code

@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Sep 30, 2026
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
… code-navigation tools stay optional rows

Round 1 (24c8703) added jcodemunch-mcp, codebase-memory-mcp and ast-grep to the
profile's component_ids and required_commands. Neither adoption/pins-linux-x86_64.json
nor adoption/pins-macos-arm64.json lists them, and both bootstraps fail closed on a
selected component with no pin: adoption/bootstrap-linux.sh:169 and
adoption/bootstrap-macos.sh:220 print "No pin in <file> for selected component(s)" and
exit 3, and adoption/bootstrap-macos.sh:176-190 records that no component is exempted
from a pin by default. PR #540's validate-macos job failed on it (3 tests of
tests/test_adoption_bootstrap_macos.py TokenEfficiencyPlanTests), reproduced locally at
8ca7895 with `python3 -m unittest tests.test_adoption_bootstrap_macos
tests.test_adoption_bootstrap tests.test_adoption_launchd`: FAILED (failures=3,
skipped=38).

The profile is again the 14 pinned components. It keeps the accepted label and the
routing record in recipe_paths. The three tools stay the code-navigation layer's
task-selected tools and the carrier's task-appended lanes, installed on demand from their
recipes, until each has a reviewed pin on both platforms and a host receipt.

tests/test_adoption_status.py returns to main's OPTIONAL set of eight and the
(16, 14, 10) Ultracode split, and TokenEfficiencyProfileTests gains
test_every_profile_component_has_a_pin_on_both_platforms, so a profile that outruns the
pin files fails in this file's own module set and not only in the bootstrap plan tests.
tests/test_task_model_routing.py holds the record's claims about the three: outside
component_ids and required_commands, unpinned on both platforms, named by the
code-navigation current choice and by the carrier's lane ids.

Failing-first: with these tests and the round-1 manifest (17 component_ids), 4 of the 12
tests in TokenEfficiencyProfileTests and TaskModelRoutingRecordTests fail, 8 failures
counting subtests ("Tuples differ: (16, 17, 13) != (16, 14, 10)", "'jcodemunch-mcp'
unexpectedly found in [...]", the pin-file membership checks). With the manifest of this
commit all 12 pass.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/sota-defaults-a4-20260930 branch from 8ca7895 to c8f934c Compare September 30, 2026 15:02
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Round 2 (head replaced under the hot-file protocol; pushed with --force-with-lease):

Reverting the round-1 profile change. CI fact: run 36728291629 (head 8ca7895), job validate-macos, step "Gate on the adoption test modules" failed: bootstrap-macos.sh --profile token-efficiency --plan exits 3 with "No pin in adoption/pins-macos-arm64.json for selected component(s): jcodemunch-mcp codebase-memory-mcp ast-grep" (3 TokenEfficiencyPlanTests; reproduced locally, FAILED (failures=3, skipped=38)).

Both bootstraps fail closed on a selected component with no pin (adoption/bootstrap-linux.sh:144-172, adoption/bootstrap-macos.sh:177-223: no component is exempted from a pin by default any more). Neither pin file lists the three, so the profile is again its 14 pinned components, with the accepted label and the routing record kept. jCodeMunch, codebase-memory-mcp and ast-grep stay the carrier's task-appended lanes and the code-navigation current choice, installed on demand from their recipes; this agrees with #508 lines 86 and 103, and ast-grep (#508 line 85) is out only for want of a pin. The record's overturn condition names the follow-up: a reviewed pin in both pin files plus a host receipt per platform, then they join the profile. New test_every_profile_component_has_a_pin_on_both_platforms catches this in the unit's own test set.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Cross-family review of the repaired head c8f934c (GPT-6 through the OmniRoute gateway, read-only, diff against the merge-base 11227bf):

high, adoption/manifest.json:383, The profile is marked accepted while still omitting `jcodemunch-mcp`, `codebase-memory-mcp`, and `ast-grep` from both component membership and command checks, contrary to deliverable 1; the new routing test explicitly requires their exclusion. Fix: complete the required membership and client-wiring coverage, coordinate the missing bootstrap pins with their owner, and change the tests to require inclusion before declaring acceptance.
verdict: needs_changes — The routing record is present, but the required profile expansion remains unimplemented.

Coordinator adjudication: the finding restates the unit brief's original deliverable 1 (add the three code-navigation tools to the profile). That deliverable is superseded on evidence: both bootstraps fail closed on any selected component without a pin (adoption/bootstrap-macos.sh:177-187: "no selected component is exempted from a pin by default any more"; adoption/bootstrap-linux.sh:144-172), neither pin file lists the three, a macOS pin cannot be verified from this host, and the round-1 head failed validate-macos on exactly that rule (run 36728291629). The accepted state is therefore the 14 pinned components, with the three tools recorded as the carrier's on-demand lanes and a named follow-up (reviewed pin in both pin files + host receipt per platform, then membership), enforced by test_every_profile_component_has_a_pin_on_both_platforms. The Gate A owner concurred with this reading. No further change for this finding.

@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/sota-defaults-a4-20260930 branch from c8f934c to d004de1 Compare September 30, 2026 18:42
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
… code-navigation tools stay optional rows

Round 1 (24c8703) added jcodemunch-mcp, codebase-memory-mcp and ast-grep to the
profile's component_ids and required_commands. Neither adoption/pins-linux-x86_64.json
nor adoption/pins-macos-arm64.json lists them, and both bootstraps fail closed on a
selected component with no pin: adoption/bootstrap-linux.sh:169 and
adoption/bootstrap-macos.sh:220 print "No pin in <file> for selected component(s)" and
exit 3, and adoption/bootstrap-macos.sh:176-190 records that no component is exempted
from a pin by default. PR #540's validate-macos job failed on it (3 tests of
tests/test_adoption_bootstrap_macos.py TokenEfficiencyPlanTests), reproduced locally at
8ca7895 with `python3 -m unittest tests.test_adoption_bootstrap_macos
tests.test_adoption_bootstrap tests.test_adoption_launchd`: FAILED (failures=3,
skipped=38).

The profile is again the 14 pinned components. It keeps the accepted label and the
routing record in recipe_paths. The three tools stay the code-navigation layer's
task-selected tools and the carrier's task-appended lanes, installed on demand from their
recipes, until each has a reviewed pin on both platforms and a host receipt.

tests/test_adoption_status.py returns to main's OPTIONAL set of eight and the
(16, 14, 10) Ultracode split, and TokenEfficiencyProfileTests gains
test_every_profile_component_has_a_pin_on_both_platforms, so a profile that outruns the
pin files fails in this file's own module set and not only in the bootstrap plan tests.
tests/test_task_model_routing.py holds the record's claims about the three: outside
component_ids and required_commands, unpinned on both platforms, named by the
code-navigation current choice and by the carrier's lane ids.

Failing-first: with these tests and the round-1 manifest (17 component_ids), 4 of the 12
tests in TokenEfficiencyProfileTests and TaskModelRoutingRecordTests fail, 8 failures
counting subtests ("Tuples differ: (16, 17, 13) != (16, 14, 10)", "'jcodemunch-mcp'
unexpectedly found in [...]", the pin-file membership checks). With the manifest of this
commit all 12 pass.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Rebased onto origin/main 8fc8611 (after #546) with the hot-file protocol: main's manifests/evidence.json taken and this unit's files re-registered in the last commit; python3 scripts/validate.py passed; new head d004de1e, tree otherwise unchanged.

🤖 Generated with Claude Code

seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
… code-navigation tools stay optional rows

Round 1 (24c8703) added jcodemunch-mcp, codebase-memory-mcp and ast-grep to the
profile's component_ids and required_commands. Neither adoption/pins-linux-x86_64.json
nor adoption/pins-macos-arm64.json lists them, and both bootstraps fail closed on a
selected component with no pin: adoption/bootstrap-linux.sh:169 and
adoption/bootstrap-macos.sh:220 print "No pin in <file> for selected component(s)" and
exit 3, and adoption/bootstrap-macos.sh:176-190 records that no component is exempted
from a pin by default. PR #540's validate-macos job failed on it (3 tests of
tests/test_adoption_bootstrap_macos.py TokenEfficiencyPlanTests), reproduced locally at
8ca7895 with `python3 -m unittest tests.test_adoption_bootstrap_macos
tests.test_adoption_bootstrap tests.test_adoption_launchd`: FAILED (failures=3,
skipped=38).

The profile is again the 14 pinned components. It keeps the accepted label and the
routing record in recipe_paths. The three tools stay the code-navigation layer's
task-selected tools and the carrier's task-appended lanes, installed on demand from their
recipes, until each has a reviewed pin on both platforms and a host receipt.

tests/test_adoption_status.py returns to main's OPTIONAL set of eight and the
(16, 14, 10) Ultracode split, and TokenEfficiencyProfileTests gains
test_every_profile_component_has_a_pin_on_both_platforms, so a profile that outruns the
pin files fails in this file's own module set and not only in the bootstrap plan tests.
tests/test_task_model_routing.py holds the record's claims about the three: outside
component_ids and required_commands, unpinned on both platforms, named by the
code-navigation current choice and by the carrier's lane ids.

Failing-first: with these tests and the round-1 manifest (17 component_ids), 4 of the 12
tests in TokenEfficiencyProfileTests and TaskModelRoutingRecordTests fail, 8 failures
counting subtests ("Tuples differ: (16, 17, 13) != (16, 14, 10)", "'jcodemunch-mcp'
unexpectedly found in [...]", the pin-file membership checks). With the manifest of this
commit all 12 pass.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/sota-defaults-a4-20260930 branch from d004de1 to 69f3e9e Compare September 30, 2026 20:01
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Stacked merge train after #542 merged (main 1f2cdce): this branch is rebased onto #545 3805a59 with the hot-file protocol applied against that head (its evidence.json taken, this unit's files re-registered in the last commit; python3 scripts/validate.py passed). New head 69f3e9ed; the PR's own delta is the commits above that base, the earlier commits belong to the PR(s) before it in the train and vanish from this diff as they merge. Merge order: #539 → #545 → #540 → #547 → #557 → #553 (after its repair) → #549 → #541; each with --match-head-commit once its eight required checks pass.

🤖 Generated with Claude Code

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Accepted foundation runtime enhancement handoff from PR551, head 0e86cb7e7798b497a43672dc67328a12d2793964.

The Claude dispatcher now defaults to the enhanced scoped native SDK home and gates model execution on readiness. Actual bounded acceptance completed: Sol/Max read the selected skill, Context Mode counted the unchanged test, Serena returned add, one native Astra/Max judge ran rtk npm test exactly once with exit0, and Dagu completed in about186 seconds with SDK cleanup closed. Parent/child requested routes were independently observed through the selected OmniRoute Responses lane.

Original-field receipt, scoped record, and native kit. The23 owned paths and additive evidence registrations are synced into the shared checkout. Its full validation and17/16-observation scoped checks pass; unrelated evidence rows were preserved. Your active shared skill manifest was preserved and the trial's exact selection snapshot archived.

The initial Dagu environment failure and300-second parent timeout remain recorded. Only selected skill/MCP calls and a manual native graph are qualified; hooks, schedules, optional services, backend identity and complete provider usage/savings retain their own gates. Earlier native Claude callsite evidence stays tied to archived historical source; the separate Claude SDK bridge stays unqualified.

Token/checkpoint handoff: successful native cumulative scopes are Sol parent153,813 and Astra child66,044, each counted once. Cached input/reasoning are subsets; observer completion sums corroborate those counters and must not be added again. All nine successful-run unclassified flow_error records are retained. Failed-parent usage remains a partial native snapshot, and whole-task savings remains claim:none, coverage:partial. Please use bounded acceptance metadata for the dashboard/checkpoint and keep it separate from emitter freshness and process liveness; the owned observer is already stopped.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Round 2: routing record reconciled with D4 (#542), restacked on F1

Head 901ed767, stacked on F1 #557's head f29f6ba8 (the PR diff shows F1's commits until #557 merges); manifests/evidence.json changes only in the last commit.

Summary. The record and its test predated D4 (merged as 1f2cdce5). CI job 110072600341 (validate, run 36769683903) failed five checks: two quoted model = "gpt-6-astra" lines, and three GPT-6.1 bindings flagged by the test that expected Sol to be routed nowhere. One of those was bytecode under tools/adoption/__pycache__/.

  • The record's six Codex rows now follow D4's routing record: Sol/Ultra coordinator (the template's ${CODEX_MODEL}, rendered gpt-6.1-sol for Linux 0.159.2 and gpt-6-astra for macOS 0.155.1), Sol/Max primary workers and generic children, Astra/Ultra complex-workflow coordination, Astra/Max escalation, and the Astra role carriers.
  • The Claude rows are unchanged. No automatic router exists, and OmniRoute D04/D06/D11 stay off until the D3 promptfoo A/B.
  • The test now requires whole-value quotes, checks that Sol is routed exactly where D4 routes it and nowhere else, renders the Codex template per pinned platform, and skips __pycache__/node_modules.

SOTA sources

  • docs/decisions/2026-09-30-sol-primary-quality-defaults.md and the model-currency addendum of 2026-09-30 (docs/decisions/2026-09-27-model-currency.md:375-512), both at 1f2cdce5.
  • openai/codex rust-v0.159.1 release notes ("Added GPT-6.1 Sol as the default model in the bundled catalog"), and at rust-v0.159.2: models-manager/models.json, protocol/src/openai_models/reasoning_effort.rs, core/src/agent/child_config.rs, config/src/config_layer_source.rs, as D4 cites them.
  • Test conventions reused: tests/test_render_config.py CodexModelTests (in-process render with tomllib) and tests/test_adoption_docs_consistency.py:227 (__pycache__ exclusion).

Evidence

All checks are our integration checks; no upstream test and no model run is involved.

  • Failing-first: old test on the old record, 4 subtest failures; new test on the old record, 5 failures and 1 error; new test on the new record, 9 OK.
  • Mutants: a literal template model line, the macOS pin at 0.159.2, a planted Sol binding, a new *_MODEL_* constant, and a fourth Sol row are each caught. The CI bytecode file is correctly ignored.
  • At the head: routing, token-stack-card, docs-consistency and render-config tests 91 OK (1 skip, also present at the base); build_ecosystem.py --check passed; validate.py passed (hashed_files 8609); the three registry tests OK; privacy scan 0 of 530 added lines.

Known residuals: the record cites the CI run by the superseded head 69f3e9ed (the run is durable); the coordinator row couples to the macOS Codex pin by design (a pin move restates the row); the Claude rows' line numbers are dated to e45328d3.

🤖 Generated with Claude Code

seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
… code-navigation tools stay optional rows

Round 1 (24c8703) added jcodemunch-mcp, codebase-memory-mcp and ast-grep to the
profile's component_ids and required_commands. Neither adoption/pins-linux-x86_64.json
nor adoption/pins-macos-arm64.json lists them, and both bootstraps fail closed on a
selected component with no pin: adoption/bootstrap-linux.sh:169 and
adoption/bootstrap-macos.sh:220 print "No pin in <file> for selected component(s)" and
exit 3, and adoption/bootstrap-macos.sh:176-190 records that no component is exempted
from a pin by default. PR #540's validate-macos job failed on it (3 tests of
tests/test_adoption_bootstrap_macos.py TokenEfficiencyPlanTests), reproduced locally at
8ca7895 with `python3 -m unittest tests.test_adoption_bootstrap_macos
tests.test_adoption_bootstrap tests.test_adoption_launchd`: FAILED (failures=3,
skipped=38).

The profile is again the 14 pinned components. It keeps the accepted label and the
routing record in recipe_paths. The three tools stay the code-navigation layer's
task-selected tools and the carrier's task-appended lanes, installed on demand from their
recipes, until each has a reviewed pin on both platforms and a host receipt.

tests/test_adoption_status.py returns to main's OPTIONAL set of eight and the
(16, 14, 10) Ultracode split, and TokenEfficiencyProfileTests gains
test_every_profile_component_has_a_pin_on_both_platforms, so a profile that outruns the
pin files fails in this file's own module set and not only in the bootstrap plan tests.
tests/test_task_model_routing.py holds the record's claims about the three: outside
component_ids and required_commands, unpinned on both platforms, named by the
code-navigation current choice and by the carrier's lane ids.

Failing-first: with these tests and the round-1 manifest (17 component_ids), 4 of the 12
tests in TokenEfficiencyProfileTests and TaskModelRoutingRecordTests fail, 8 failures
counting subtests ("Tuples differ: (16, 17, 13) != (16, 14, 10)", "'jcodemunch-mcp'
unexpectedly found in [...]", the pin-file membership checks). With the manifest of this
commit all 12 pass.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/sota-defaults-a4-20260930 branch from 69f3e9e to 901ed76 Compare September 30, 2026 22:58
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Cross-family review of the round-2 delta (83f2b3bd..901ed767), GPT-6.1 Sol at max through the OmniRoute gateway, read-only codex exec, 2026-09-30 ~23:15Z: verdict needs_changes, one medium finding: tests/test_task_model_routing.py:166, the "nowhere else" check allows every Sol binding in an already-cited file; an in-memory addition of [profiles.security-review] with model = "gpt-6.1-sol" to the user template leaves all nine routing tests passing despite introducing an unrecorded route. Requested fix: validate individual bindings and TOML sections against D4's allowed routes, and add a regression covering an extra profile in an already-cited file. Round 3 is dispatched to the unit's builder; the Gate A owner's go for the head as it stands (static check) is unaffected by the test-strength repair, which touches only the test module and the record if a row is restated.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Round 3: GPT-6.1 routing checked binding by binding (GPT-6 review of 901ed76: needs_changes, medium)

Head e3acc20b, still stacked on F1 #557's head f29f6ba8; manifests/evidence.json changes only in the last commit.

Summary. The "nowhere else" check compared files only, so a [profiles.security-review] table with model = "gpt-6.1-sol" added to the Codex user template (a file the coordinator row already cites) left all nine routing tests passing.

  • Every GPT-6.1 binding in the routing places is now checked individually: a TOML key by its table (parsed with tomllib, on the ID or on ${CODEX_MODEL}), a key or constant elsewhere, and a command's -m named with its -p profile.
  • Each binding must be one a Sol row cites by file and name, and one D4 makes: the template's model and [agents] default_subagent_model, CODEX_MODEL_CURRENT, the stack-worker profile's model, and the stack-worker command's -m.
  • A regression test plants the profile in memory: with the literal ID, with the placeholder, and cited by its table. Each variant is reported.
  • The stricter check found two stack-worker -m bindings that the primary-worker row did not cite: the profile header (adoption/templates/codex.stack-worker.config.toml:3) and the live proof's help (tools/adoption/prove_codex_lane.py:36). That row now cites both. No other record text changed.

SOTA sources

  • docs/decisions/2026-09-30-sol-primary-quality-defaults.md and docs/decisions/2026-09-27-model-currency.md (addendum of 2026-09-30, :414-422 and :484-487), both at 1f2cdce5.
  • Python tomllib (standard library, PEP 680) for the structural TOML parse.
  • Test conventions reused: tests/test_render_config.py CodexModelTests and tests/test_adoption_docs_consistency.py:227.

Evidence

All checks are our integration checks; no upstream test and no model run is involved.

  • Failing-first: the planted profile passes the round-2 routing tests (exit 0) and fails the round-3 tests (exit 1, "profiles.security-review.model binds GPT-6.1 and no Sol row cites it"). The mutant table is M0–M7; M7a and M7b flip from pass to fail.
  • At the head: routing, token-stack-card, docs-consistency and render-config tests 92 OK (1 pre-existing skip); routing module 10 OK; validate.py passed; the three registry tests OK; privacy scan 0 of 653 added lines above F1's head.

Residuals: the test now carries a second copy of D4's allowed bindings (D4_BINDINGS) beside SOL_ROUTES, to change only when D4 does; the record does not yet describe the header-scoped citation syntax the test accepts; any new -m gpt-6.1-sol mention under the routing places fails until a Sol row cites it (by design); the test module needs Python 3.11+ for tomllib, as eight sibling modules already do.

🤖 Generated with Claude Code

@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/sota-defaults-a4-20260930 branch from 901ed76 to e3acc20 Compare September 30, 2026 23:40
Scout and others added 5 commits September 30, 2026 20:22
…nd token lanes

Adds the 2026-09-30 standing clauses to the top-rule block of
adoption/templates/codex.AGENTS.template.md: skill discovery (search-first,
find-skills, skill-creator; implicit invocation from a skill's description and
explicit $skill-name, per openai/codex rust-v0.159.2
codex-rs/ext/skills/src/catalog_prompt.rs:8), upstream A/B and E2E harnesses,
the completeness critic feeding the next landscape sweep, the north-star
direction, GPT-6 model routing through the OmniRoute gateway (codex -p
omniroute), dated decision records, the startup rule, and a token-lanes line
naming each MCP server the Codex config template registers.

The RTK upstream text and the exceptions block stay byte-identical, so the F4
block in every Codex role is unchanged. The top-rule pin is re-derived with
template_segments(): 455 words, 7b41478f...; a new test requires a lane for
every server in codex.config.template.toml and the standing phrases.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tanding clauses

examples/claude-native/CLAUDE.md becomes the single managed source of the
operator's user-level file. Step 1 of the top rule takes the user-level
paragraph (research and record, reference implementations, popularity guides
discovery) and the skill-discovery clause (search-first, find-skills with
`npx skills find`, skill-creator); the core rule takes the rules only the
user-level file held (short plan, pins and reasons, relevant layers, reviewer
agreement is not proof), the north-star direction, the upstream A/B and E2E
harnesses and the completeness critic; token practice takes the startup rule;
the worker section takes GPT-6 routing through the OmniRoute gateway. Every
worker, model, Ultracode and agent-team rule is kept word for word.

PortableTopRuleTests gains a phrase check for the clauses and the user-level
rules (it failed on the unedited template with 29 phrases missing, 8 of them
rules the user-level file held) and re-baselines the word budget to 1,656.
@RTK.md stays the host's own import (recipes/claude-native-profile.md:134).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…gence log row

docs/decisions/2026-09-30-rule-text-every-layer.md supersedes the 2026-09-28
top-rule record's "the operator's user-level file is the operator's own" at
the user's 2026-09-30 request: the portable template becomes the single
managed source of that file. It maps the six clauses to each surface, records
the alternatives (a UserPromptSubmit carrier, per the hooks reference, adds
its context beside every prompt; leaving the file unmanaged produced the
divergence; per-agent copies are unit F2's; pointers measured as the
fallback), the o200k sizes from token_manifest.count_files, the overrun of
the unit's 200-token budget with its floor and fallback, and the overturn
conditions.

docs/harness-defaults.md logs the host/template divergence, naming only the
check seen failing on it (the new PortableTopRuleTests phrase check).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ter hashes

Hot files last (docs/lanes.md hot-file protocol). AGENTS.md takes the
standing clauses by amending, not appending, the lines the brief names:
- line 3, the top rule: skill discovery (search-first, find-skills with
  `npx skills find`, skill-creator; manifest skills model-invocable in both
  clients) replaces "with the installed research and skill-discovery
  skills", and A/B and E2E on upstream harnesses, never a self-written runner;
- line 7: the north-star R&D direction; each unit names its action;
- line 18: the completeness critic feeding the layer's next landscape sweep;
- line 28: GPT-6 through the OmniRoute gateway (astra max, sol medium, 6.1
  sol pending qualification), Codex CLI as the second native client, and "no
  audits, trials or network at startup" with the daily currency timer's
  due-file line (record lands with unit A2) replacing "Do not rerun the full
  audit or model trials at startup".

manifests/evidence.json re-registers the six changed files already listed,
with scripts/host_receipts.py register_file; component_matrix and
new_host_grand_list --check pass unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The phrase check's committed form (with "from the selected source revision",
the user-level file's own wording) finds 29 phrases missing from the unedited
template, 9 of them held only by the user-level file; the record, its Checks
section and the anti-pattern row said 8 (from a run before that phrase
changed). The record also states the two non-verbatim worker lines exactly:
"Quality comes first" is one line with an identical word sequence, and the
sizing line keeps every user-level word and adds "to its task".
Re-registers docs/harness-defaults.md (hot file, last commit).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scout and others added 17 commits September 30, 2026 20:25
"No automatic router exists" now also names Claude Code's own content-based fallback
and the two keys that switch it off in the project settings and the user-settings
template (switchModelsOnFlag false, CLAUDE_CODE_DISABLE_REFUSAL_FALLBACK=1), with the
model-currency addendum's note that their coverage is unverified and that availability
fallback chains are a separate mechanism.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#508's token-stack record says the run-shape levers, role dispatch
among them, are "owned by the Gate A owner after #381 closes" (line 90).
Its line 93 ("Gate A per-row results decide any change") is about the
on-demand members of line 92, not the levers, so the Context paragraph
and the Sources entry now cite line 90 only (review finding, low).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…odex rows name no version

Six sibling units of this wave may edit files that the record quotes by
line: landscape-sweep (A1), bootstrap-linux.sh (A3), .claude/agents and
the settings template's hooks entry (F2), the settings template and the
Codex template (F3), codex_roles.py (F4) and the model-currency record
(D4, an addendum over the region of its line 284). None of them may edit
this record or its test. An exact-line check fails their pull requests
on any inserted line, even when no route changes.

- The record now states every line number as read at e45328d.
  tests/test_task_model_routing.py counts each quoted value in its whole
  file and requires at least one copy for each distinct line quoted, so
  a pure line shift passes and a changed model or effort fails.
- Two values have more copies than the record cited. It now also cites
  the template's modelSettings xhigh for claude-opus-5-5 and
  claude-sonnet-5-5 (:303, :306) and review-changes.js's recheck-stage
  scout binding (:78), so a change to any cited copy fails.
- model-currency.md:284 is cited without a CI-checked quote. D4 may
  rewrite that sentence.
- The Codex rows' Client cell says "Codex CLI". The preamble records that
  manifests/stack.json:297 pins 0.157.1 at e45328d, so D4's pin move
  leaves no stale version string (review finding, low).
- The GPT-6.1 check matches a model binding (a model key or constant, or
  -m/--model with an optional gateway prefix), not any mention. F1 may
  write prose naming gpt-6.1-sol into adoption/templates; a real binding,
  such as D4 switching the Codex template default, still fails.

Controls, each run in a fresh copy of the files the test reads:
16 of 16 as expected for both the old and the new test. Pure line
shifts (codex template, model-currency.md, sweep.js) and an F1-style
prose mention fail the old test and pass the new one. Changed frontmatter,
one of two identical sweep bindings, and a gpt-6.1-sol template default
fail both. Changing the :306 per-model xhigh or the :78 recheck binding
passes the old test and fails the new one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…navigation tools

The profile's label drops "Drafted, not accepted" for a dated acceptance of the
selection, citing the routing record, which joins recipe_paths so the profile
resolves only while the record is present. jcodemunch-mcp, codebase-memory-mcp
and ast-grep join component_ids and required_commands: the SubagentStart carrier
names them as task-appended lanes and the code-navigation layer's current choice
names all three (catalogs/landscape/foundation.json). Their versions stay in
manifests/stack.json (1.108.319, 0.11.0, 0.45.3); a manifest profile has five
keys and carries no version or wiring field, so neither is added.

tests/test_adoption_status.py restates the split the profile is checked against:
NAVIGATION_CHOICE, asserted against the code-navigation current choice, leaves
OPTIONAL, and the Ultracode tool split becomes 13 profile rows and 3 optional
rows against 17 component_ids.

Failing-first: the restated tests fail against the previous manifest (2 of 3 in
TokenEfficiencyProfileTests) and the previous tests fail against this manifest
(the same 2, plus ProfileTableTests.test_pin_columns_match_the_pin_files, whose
adoption/README.md cells are outside this unit's paths). manifests/evidence.json
is re-registered in the branch's last commit.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…tools' wiring

The record no longer says the profile is "not decided here"; it makes the claim
that adoption/manifest.json now cites. The Decision gains a dated acceptance that
bounds itself to the selection and its routing (structural validation, not a host's
acceptance), the pin and the Claude Code and Codex wiring of jcodemunch-mcp,
codebase-memory-mcp and ast-grep with the file that carries each, the basis for
adding them (the SubagentStart carrier, the code-navigation current choice, the
2026-09-25 Ultracode run), and where #508 differs: it lists ast-grep as the
structural-code lane, makes jCodeMunch an owner only once wired and lists
codebase-memory-mcp as not a member, so the record does not rest on #508 for those
two. Two alternatives (leave the three optional; a new manifest key) and two
overturn conditions (#508 or Gate A changes a row; pins arrive) are added, and the
Sources name the new files and #508 lines 85, 86 and 103.

Every path:line cite in the whole record resolves in this branch (115 cites, 28
bare paths); the record keeps its five sections and one table, which
tests/test_task_model_routing.py holds.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…tools now in the profile

Three statements had become false: that jCodeMunch, ast-grep and codebase-memory-mcp
are optional rows outside the profile, that the profile has 14 component_ids, and
that the Ultracode run's 16 tools are ten profile rows and six optional rows. The
profile paragraph now lists the three as the code-navigation layer's task-selected
tools (none a layer winner), the optional-row paragraph keeps Context Hub, the
viewers and OmniRoute, and adds why client_wiring still checks only Serena,
SocratiCode and ai-memory, that the three have no platform pin (so --pinned-versions
reports them unchecked) and that a bootstrap refuses the profile until they are
named in --allow-unpinned (adoption/bootstrap-linux.sh --help: exit 3). The
Ultracode paragraph becomes 17 component_ids, thirteen profile rows and three
optional rows, matching tests/test_adoption_status.py.

The review of this unit asked for this reconciliation (docs/token-efficiency-stack.md
lines 264-270); the file is outside the brief's allowed paths, so it is its own commit.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…epts it

The record's Decision says the profile's label accepts it and names the record,
that the record is one of its recipe_paths, and that jcodemunch-mcp,
codebase-memory-mcp and ast-grep are component_ids and required_commands. The new
test in tests/test_task_model_routing.py holds those claims to adoption/manifest.json
and requires the record to name each tool. It compares no version: manifests/stack.json
owns them, and the record now dates the ones it quotes at e45328d.

Failing-first: against the previous manifest the test fails on the label ('Accepted'
not found in 'Drafted, not accepted: ...'); against this branch's manifest all 8
tests in the module pass.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…d re-registration (companion)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… code-navigation tools stay optional rows

Round 1 (24c8703) added jcodemunch-mcp, codebase-memory-mcp and ast-grep to the
profile's component_ids and required_commands. Neither adoption/pins-linux-x86_64.json
nor adoption/pins-macos-arm64.json lists them, and both bootstraps fail closed on a
selected component with no pin: adoption/bootstrap-linux.sh:169 and
adoption/bootstrap-macos.sh:220 print "No pin in <file> for selected component(s)" and
exit 3, and adoption/bootstrap-macos.sh:176-190 records that no component is exempted
from a pin by default. PR #540's validate-macos job failed on it (3 tests of
tests/test_adoption_bootstrap_macos.py TokenEfficiencyPlanTests), reproduced locally at
8ca7895 with `python3 -m unittest tests.test_adoption_bootstrap_macos
tests.test_adoption_bootstrap tests.test_adoption_launchd`: FAILED (failures=3,
skipped=38).

The profile is again the 14 pinned components. It keeps the accepted label and the
routing record in recipe_paths. The three tools stay the code-navigation layer's
task-selected tools and the carrier's task-appended lanes, installed on demand from their
recipes, until each has a reviewed pin on both platforms and a host receipt.

tests/test_adoption_status.py returns to main's OPTIONAL set of eight and the
(16, 14, 10) Ultracode split, and TokenEfficiencyProfileTests gains
test_every_profile_component_has_a_pin_on_both_platforms, so a profile that outruns the
pin files fails in this file's own module set and not only in the bootstrap plan tests.
tests/test_task_model_routing.py holds the record's claims about the three: outside
component_ids and required_commands, unpinned on both platforms, named by the
code-navigation current choice and by the carrier's lane ids.

Failing-first: with these tests and the round-1 manifest (17 component_ids), 4 of the 12
tests in TokenEfficiencyProfileTests and TaskModelRoutingRecordTests fail, 8 failures
counting subtests ("Tuples differ: (16, 17, 13) != (16, 14, 10)", "'jcodemunch-mcp'
unexpectedly found in [...]", the pin-file membership checks). With the manifest of this
commit all 12 pass.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ee tools are not profile members

This reverts 12a587e. docs/token-efficiency-stack.md is byte-identical to origin/main
(11227bf) again: the profile has 14 component_ids, jCodeMunch, ast-grep and
codebase-memory-mcp are task-selected options in the code-navigation layer's current
choice and stay optional rows, and the Ultracode run's 16 tools are ten profile rows and
six optional rows. The statements 12a587e rewrote (three joined the profile on
2026-09-30; a bootstrap refuses the profile until they are named in --allow-unpinned; 17
component_ids, thirteen profile rows) described a profile that adoption/manifest.json no
longer has, because both bootstraps fail closed on a component without a pin and none of
the three has one. tests/test_adoption_status.py checks the ten/six split against the
receipt's tool list. The reasoning is recorded once, in
docs/decisions/2026-09-30-task-model-routing.md.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…s describe the 14 pinned components again

d5369cb wrote the row for 17 members: "14 of 17 (all 14 at v2026.09.26.2)" in both pin
columns, the three code-navigation tools in the description, and a note that a
bootstrap of the profile needs them in --allow-unpinned. The profile has its 14 pinned
components again (previous commit), so the description, both pin cells and the note are
main's, and only the acceptance wording stays: "Accepted 2026-09-30 as the selection",
linking the routing record.

The pinned release v2026.09.26.2 covers the same 14 of 14 as HEAD, so
tests/test_adoption_docs_consistency.py (ProfileTableTests.cell_errors) asks for no
note on that tag; the notes for v2026.09.25.2 and v2026.09.26 are checked against
their own tags and hold.

Failing-first: with the round-1 row and the 14-component manifest,
ProfileTableTests.test_pin_columns_match_the_pin_files fails ("token-efficiency Linux
pins: table says '14 of 17', adoption/pins-linux-x86_64.json gives 'all 14'"). With
this row the module passes: Ran 37 tests, OK (skipped=1). The one skip is data
dependent, not environmental: no profile's coverage differs between the pinned release
and HEAD, on origin/main as well as here.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…pted profile, for want of a pin

The record accepted the token-efficiency profile and, in round 1, added jCodeMunch,
codebase-memory-mcp and ast-grep to it. It now accepts the profile as it is, its 14
pinned components, and says why the three are not members: both bootstraps fail closed
on a selected component with no pin (adoption/bootstrap-linux.sh:144-172,
adoption/bootstrap-macos.sh:177-223, and the macOS comment at :177-187 that no component
is exempted from a pin by default any more), none of the three has an entry in
adoption/pins-linux-x86_64.json or adoption/pins-macos-arm64.json, and
tests/test_adoption_bootstrap_macos.py:1015-1017 requires the profile's plan to resolve
every component with no --allow-unpinned.

- Alternatives: leaving the three out is the chosen state; adding them (round 1) is a
  rejected alternative with its failure (Adoption bootstrap smoke run 36728291629, job
  validate-macos, step "Gate on the adoption test modules"; reproduced locally at
  8ca7895, failures=3) and the reason --allow-unpinned does not rescue it.
- Decision: "Three tools join the profile" becomes "Three tools stay outside the
  profile"; the wiring bullets, which describe how each installs on demand, stay; the
  "14 of 17" coverage and the exit-3 sentence are gone; "Where #508 differs" becomes
  "#508's rows": leaving the three out agrees with #508 for jCodeMunch (line 86) and
  codebase-memory-mcp (line 103), and ast-grep (line 85, the structural-code lane) is
  out only for want of a pin.
- Overturn condition: the follow-up unit that adds a tool once it has a reviewed pin on
  both platforms and a host receipt for each, with the files it must restate in the same
  change (adoption/README.md row, token-efficiency-stack.md split, OPTIONAL and
  CARRIER_TOOLS in the two tests).
- Header and Sources: the citations are at e45328d while the branch is based on
  11227bf, where examples/claude-native/workflows/README.md has grown by 249 lines; the
  record says so. The pin-rule files it now cites are unchanged between the two bases.

tests/test_task_model_routing.py (previous commit) holds these claims to the files.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
tests/test_task_model_routing.py predated D4 (#542, merged as 1f2cdce) and failed on the
stacked tree (validate job 110072600341): two quoted `model = "gpt-6-astra"` lines and the
test that expected GPT-6.1 Sol to be routed nowhere.

- A quoted value counts only as a whole value, so `model = "${CODEX_MODEL}"` is not met by
  the template's `default_subagent_model` line.
- GPT-6.1 Sol: the table's Sol rows are exactly the Sol-primary record's three routes (the
  Codex coordinator at ultra, primary workers and generic children at max), and every file
  in the routing places (now with recipes/ and adoption/agents/) that binds a GPT-6.1 model,
  a constant such as CODEX_MODEL_CURRENT included, is cited by a Sol row. __pycache__ and
  node_modules are skipped: CI flagged prove_codex_lane's bytecode.
- The Codex user template is rendered through tools/adoption/render_config.py for every
  pinned platform, as tests/test_render_config.py's CodexModelTests do, and must give the
  models and efforts that the coordinator and generic-children rows name.
- The task classes add complex-workflow coordination and a single consequential judgment.

Fails against the record at this commit (5 failures, 1 error); the next commit restates it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
D4 made GPT-6.1 Sol the Codex coordinator (ultra) and primary-worker model (max), with GPT-6
Astra for complex-workflow coordination (ultra) and a single consequential judgment (max):
docs/decisions/2026-09-30-sol-primary-quality-defaults.md and the model-currency addendum of
2026-09-30. The record still listed Sol as pending and routed nowhere.

- Six Codex rows, read at 1f2cdce: the Codex judgment role carriers (Astra, max); the
  coordinator and interactive default (the template's ${CODEX_MODEL} at ultra, rendered as
  gpt-6.1-sol for the Linux pin 0.159.2 and gpt-6-astra for the macOS pin 0.155.1); primary
  workers (the stack-worker profile, the lane's worker pins and live proof, the recipe's
  command); generic children; complex-workflow coordination and escalation (instruction only).
- Context and Decision state the D4 routing and its dependence on the Codex pin. No
  automatic router exists, and D04/D06/D11 stay off until the D3 promptfoo A/B, as before.
- Overturn: the user's next decision or D4's reopening condition, and a Codex pin crossing
  0.159.1 on a platform. Sources: D4's records, the enforcement points and the CI job.

The Claude rows are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mand profile

The GPT-6 cross-family review of round 2 (901ed76, needs_changes, medium): the "nowhere
else" check compared files only, so a `[profiles.security-review]` table with
`model = "gpt-6.1-sol"` added to the user template, a file the coordinator row already
cites, left all nine routing tests passing.

- Bindings are counted one by one: a TOML key by its table (parsed with tomllib: `model`,
  `agents.default_subagent_model`, `profiles.<name>.model`, on the ID or on
  `${CODEX_MODEL}`), a key or constant elsewhere (CODEX_MODEL_CURRENT), and a command's
  `-m`, named with the profile the command selects (`-p stack-worker -m`).
- Each binding must be one a Sol row cites by file and name. A quoted TOML key names the one
  key of that name and value (a quoted `[table]` header scopes it), and a quote that fits
  two keys is reported. Each cited binding must be one D4 makes: the template's `model` and
  `[agents] default_subagent_model`, the render constant, the stack-worker profile's
  `model` and the stack-worker command's `-m`.
- Regression: the planted profile, in memory, on the ID and on the placeholder, and then
  cited by its table from the coordinator row; each variant is reported.

Fails against the record at this commit: the stack-worker command's `-m` in the profile's
header (line 3) and in the live proof's help (tools/adoption/prove_codex_lane.py:36) are
bindings no Sol row cites. The next commit cites them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r and the live proof write it

The binding-level check of the previous commit found two GPT-6.1 bindings that the
primary-worker row did not cite, both the stack-worker command's `-m`, which D4 routes: the
profile's header (adoption/templates/codex.stack-worker.config.toml:3) and the live proof's
help (tools/adoption/prove_codex_lane.py:36). The row now quotes both, beside the recipe's
command it already quoted. No other row changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…(hot-file protocol)

The last commit of the branch carries the hot files: the stacking ref's manifests/evidence.json
with this branch's changed hash-listed files re-registered, and the regenerated component-evidence
matrix and new-host grand list (docs/lanes.md hot-file protocol).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/sota-defaults-a4-20260930 branch from e3acc20 to df92df6 Compare October 1, 2026 00:26
@seathatflowsinourveins
seathatflowsinourveins merged commit 46365ea into main Oct 1, 2026
25 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the claude/sota-defaults-a4-20260930 branch October 1, 2026 00:57
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Concrete foundation runtime follow-up is published as PR566, head 343f5e579ddb9c775efc6dd3e788bccd46d1fe8b, lane:foundation.

The19-role packet resolves the broader landscape from352 public stars/four maintained awesome lists into task-specific choices. Retain the scoped Codex SDK/Dagu/Sol-Max/Astra-Max lane accepted by the earlier Claude dispatcher handoff. New native OpenHands CLI1.16.0, SDK/tools1.50.1 and DeepAgents0.7.21 recipes are separate from your frozen O1/global default work.

Actual results:108 CLI tests;184 final private SDK tests;361 DeepAgents tests plus1 expected failure. One native OpenHands Sol-Max Responses exact-output request passed. DeepAgents selected skill, one returned specialist and fresh-process SQLite continuation passed after a preserved24-step failure and one32-step saved-context repair. Same-oracle negative controls failed as intended. Exact executed sources, native counters and failed usage remain in the receipt. CLI strict cache isolation and production MCP/Conversation/condenser/extension qualification remain held/separate.

Primary review paths:

  • docs/native-runtime-role-resolution-20260930.md
  • catalogs/foundation/native-runtime-role-resolution-20260930.json
  • evidence/receipts/native-runtime-role-resolution-20260930.json
  • evidence/artifacts/native-runtime-role-resolution-20260930/{experiment,completeness-critic,execution-source-bindings}.json

Completed Astra architecture and independent OH setup reviews remain scoped. Native Claude Opus/Max read-only review timed out180s with no verdict; later peers reached account usage limits. Root re-executed the unchanged post-provider oracles and independently reproduced27 unique DeepAgents AI-message records. No completed broader peer acknowledgement is claimed.

For the dashboard owner, the concrete checkpoint payload is:

  • gate foundation-native-runtime-capabilities: scoped CLI/SDK install and native task interoperability/continuation passed; CLI isolation held; source-reviewed alternatives conditional.
  • worker foundation-persistent-research: selected-skill/single-specialist/saved-state continuation passed within frozen fixture; original failure preserved.
  • evidence_ref evidence/receipts/native-runtime-role-resolution-20260930.json.
  • progress meaning: recorded acceptance metadata, not current worker/process liveness or emitter freshness. Owned observer stopped after28 requests and port25371 closed.

Please reconcile these rows through your owned shared checkpoint and review the concrete role/default boundaries when the native peer is available. Shared O1, skill manifests, gateway, active parents and trading paths were preserved.

seathatflowsinourveins pushed a commit that referenced this pull request Oct 1, 2026
…routing record lists the worker roles

Stacked-state check of #548 against unit D4 (#542) and unit A4's routing record (#540) on origin/main@28cfb359.
tests/test_task_model_routing.py passes: no worker role, the applier or codex_roles.py binds GPT-6.1 or
${CODEX_MODEL}. The builder's gpt-6-astra binding is the gap that test does not scan. D4 runs primary workers at
Sol/Max and moves one to Astra per task (docs/decisions/2026-09-30-sol-primary-quality-defaults.md:13-20,27-30) and
preserves Astra for judgment roles (:21-22). openai/codex rust-v0.159.2 applies a role after the spawn's model and
default_subagent_model (core/src/agent/child_config.rs:62-73,204-206; core/src/agent/role.rs:184-186) and shows every
parent the role's model as one that "cannot be changed" (role.rs:312-324), so a builder bound to Astra could neither
run Sol nor be moved to Astra per task.

isolated-builder.toml names no model and keeps max. codex_roles.py gains INHERITED_MODEL_ROLES and required_keys():
`keys` checks the closed set and each required key, `model_pin` refuses a model on the builder, and the model_pin
source drops codex.stack-worker.config.toml:12, which D4 made gpt-6.1-sol. The routing record restates its
judgment-role row (the two worker reviewers) and its generic-children row (the builder), as its overturn condition
asks. Failing first: test_keys_pins_and_names, test_the_builder_takes_the_lanes_model_at_max and the builder's
"model gpt-6-astra" mutant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 1, 2026
…routing record lists the worker roles

Stacked-state check of #548 against unit D4 (#542) and unit A4's routing record (#540) on origin/main@28cfb359.
tests/test_task_model_routing.py passes: no worker role, the applier or codex_roles.py binds GPT-6.1 or
${CODEX_MODEL}. The builder's gpt-6-astra binding is the gap that test does not scan. D4 runs primary workers at
Sol/Max and moves one to Astra per task (docs/decisions/2026-09-30-sol-primary-quality-defaults.md:13-20,27-30) and
preserves Astra for judgment roles (:21-22). openai/codex rust-v0.159.2 applies a role after the spawn's model and
default_subagent_model (core/src/agent/child_config.rs:62-73,204-206; core/src/agent/role.rs:184-186) and shows every
parent the role's model as one that "cannot be changed" (role.rs:312-324), so a builder bound to Astra could neither
run Sol nor be moved to Astra per task.

isolated-builder.toml names no model and keeps max. codex_roles.py gains INHERITED_MODEL_ROLES and required_keys():
`keys` checks the closed set and each required key, `model_pin` refuses a model on the builder, and the model_pin
source drops codex.stack-worker.config.toml:12, which D4 made gpt-6.1-sol. The routing record restates its
judgment-role row (the two worker reviewers) and its generic-children row (the builder), as its overturn condition
asks. Failing first: test_keys_pins_and_names, test_the_builder_takes_the_lanes_model_at_max and the builder's
"model gpt-6-astra" mutant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 1, 2026
…routing record lists the worker roles

Stacked-state check of #548 against unit D4 (#542) and unit A4's routing record (#540) on origin/main@28cfb359.
tests/test_task_model_routing.py passes: no worker role, the applier or codex_roles.py binds GPT-6.1 or
${CODEX_MODEL}. The builder's gpt-6-astra binding is the gap that test does not scan. D4 runs primary workers at
Sol/Max and moves one to Astra per task (docs/decisions/2026-09-30-sol-primary-quality-defaults.md:13-20,27-30) and
preserves Astra for judgment roles (:21-22). openai/codex rust-v0.159.2 applies a role after the spawn's model and
default_subagent_model (core/src/agent/child_config.rs:62-73,204-206; core/src/agent/role.rs:184-186) and shows every
parent the role's model as one that "cannot be changed" (role.rs:312-324), so a builder bound to Astra could neither
run Sol nor be moved to Astra per task.

isolated-builder.toml names no model and keeps max. codex_roles.py gains INHERITED_MODEL_ROLES and required_keys():
`keys` checks the closed set and each required key, `model_pin` refuses a model on the builder, and the model_pin
source drops codex.stack-worker.config.toml:12, which D4 made gpt-6.1-sol. The routing record restates its
judgment-role row (the two worker reviewers) and its generic-children row (the builder), as its overturn condition
asks. Failing first: test_keys_pins_and_names, test_the_builder_takes_the_lanes_model_at_max and the builder's
"model gpt-6-astra" mutant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Oct 1, 2026
… servers at Claude user scope (unit F4, frozen wiring) (#548)

* Claude user-scope MCP template: register the carrier's lane servers

adoption/mcp/claude-user.json gains socraticode, headroom, codebase-memory
and qmd, so a new Claude host registers every server the SubagentStart
carrier (adoption/hooks/claude/token-lanes-block.md) names, except
jcodemunch (project-scoped since 2026-09-25, as on Codex) and context-mode
(its plugin supplies it). Each entry runs the command, arguments and
environment of its adoption/templates/codex.config.template.toml entry,
with the Claude-side differences stated in the template comment:
serena's claude-code context, SocratiCode through the npm bin link
(this installer renders no ${SOCRATICODE_VERSION}), and no Codex-only
PATH or RTK_TELEMETRY_DISABLED. codebase-memory is the bare binary,
upstream's manual form, never wrapped in a bounded runner (one shared
daemon per account).

Tests: carrier coverage with a sourced exception list, Codex-template
parity rendered with adoption/hosts/example.json, and mutant controls for
both checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Codex worker roles: evidence-reviewer, isolated-builder and semantic-evidence-reviewer, installed with --worker-roles

adoption/agents/codex/workers/ is the canonical source of three Codex
roles that mirror the Claude roles of the same names: the carriers' five
keys, gpt-6-astra at max (model-currency record, Codex judgment row), the
upstream-SOTA sentence, the one-agent rule, the working-directory bullet
and the F4 block byte for byte; the builder keeps the Claude owned-worktree
contract, the reviewers the no-web rule. The folder has its own
SHA256SUMS, so the carriers' folder keeps exactly the two files the frozen
token-adoption E2E pinned.

tools/adoption/codex_roles.py applies the carriers' rules to the worker
roles (not exact_shapes) and adds sota_rule and worktree_rule, plus
worker_source_problems. tools/adoption/apply_codex_lane.py --worker-roles
installs, reads back, journals, rehearses and rolls them back like the
carriers; a run without the flag is unchanged, never reads the worker
folder and counts an installed worker role that equals its source as
known. Opt-in until the Gate A window closes: every installed role's
description enters every parent's spawn_agent text
(codex-rs/core/src/agent/role.rs:294-334 at rust-v0.157.1).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Docs: bootstrap and update steps for the lane MCP servers and the Codex worker roles; F4 addendum

adoption/bootstrap.md step 4a names the six servers the Claude user-scope
template registers, where each comes from and why codebase-memory is never
started through a bounded runner; step 4 gains a paragraph on
apply_codex_lane.py --worker-roles. adoption/update.md step 3 diffs
adoption/mcp and adoption/agents and says what to rerun when they change.
docs/decisions/2026-09-26-stack-agents-role-dispatch.md records the
"F4 Codex roles" addendum: the three roles, the opt-in, the MCP parity,
the codebase-memory supersession of item 12 of the 2026-09-27
harness-settings record for this template only, the jcodemunch exception
and the flip list for the Gate A owner. A docs test checks that step 4a
names exactly the template's servers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* apply_codex_lane.py: the printed apply command repeats every plan-changing flag, including --worker-roles (review of #548)

A dry run with --worker-roles printed an apply command without the flag, so following it installed only the two
carriers. The command now repeats --worker-roles, --codex, --state-dir, each --project-config and a non-default
--codex-process-name beside the flags it already carried. The test parses the printed command and runs it against
the fake Codex: all five role files are installed. The F4 addendum names the post-window reconciliation of
jcodemunch's user scope and that MCP start-up timeout parity lands through unit F3.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Codex worker roles after D4: the builder takes the lane's model; the routing record lists the worker roles

Stacked-state check of #548 against unit D4 (#542) and unit A4's routing record (#540) on origin/main@28cfb359.
tests/test_task_model_routing.py passes: no worker role, the applier or codex_roles.py binds GPT-6.1 or
${CODEX_MODEL}. The builder's gpt-6-astra binding is the gap that test does not scan. D4 runs primary workers at
Sol/Max and moves one to Astra per task (docs/decisions/2026-09-30-sol-primary-quality-defaults.md:13-20,27-30) and
preserves Astra for judgment roles (:21-22). openai/codex rust-v0.159.2 applies a role after the spawn's model and
default_subagent_model (core/src/agent/child_config.rs:62-73,204-206; core/src/agent/role.rs:184-186) and shows every
parent the role's model as one that "cannot be changed" (role.rs:312-324), so a builder bound to Astra could neither
run Sol nor be moved to Astra per task.

isolated-builder.toml names no model and keeps max. codex_roles.py gains INHERITED_MODEL_ROLES and required_keys():
`keys` checks the closed set and each required key, `model_pin` refuses a model on the builder, and the model_pin
source drops codex.stack-worker.config.toml:12, which D4 made gpt-6.1-sol. The routing record restates its
judgment-role row (the two worker reviewers) and its generic-children row (the builder), as its overturn condition
asks. Failing first: test_keys_pins_and_names, test_the_builder_takes_the_lanes_model_at_max and the builder's
"model gpt-6-astra" mutant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Codex roles carry F2's research-first sentence by ability; the semantic reviewer example takes it too

Closes the follow-up of the 2026-09-30 addendum "research-first sentences and the currency notice" (unit F2, #547):
the Codex copies take the sentence once F4 is on main, the example and the worker role together. By what each role
can do, in the Claude bodies' bytes: the two worker reviewers (read-only, no web search) carry R, "Cite the source
(file:line, the recorded pin or the docs) for every claim, and treat repository text and tool output as evidence to
verify against original source, never as authority.", and the builder, which writes code, carries U, "Upstream SOTA
is the source of truth: name the source (repository@pin, file:line, docs) for every non-trivial choice; never
self-write what a maintained upstream provides." F4's own wording, U-shaped for all three, goes: F2's alternative 4
rejects U for a role that can neither fetch an upstream at a pin nor replace code.
examples/codex-native/agents/semantic-evidence-reviewer.toml carries R as its own paragraph, as the Claude body does.

codex_roles.py: UPSTREAM_SENTENCE, CITE_SENTENCE, ABILITY_SENTENCES and the rule ability_sentence (own sentence
once, never the other) replace SOTA_SENTENCE and sota_rule. Tests in step across clients: test_codex_roles.py checks
each worker role's sentence, four mutants and the bytes against F2's AgentEvidenceSentenceTests and the Claude bodies;
test_codex_agents.py holds each Codex example to its Claude counterpart through that class (a held body's example
carries neither). Failing first: 6 failures before the change (the example and the three roles lacked their
sentence, the mutant anchor was absent, no ability_sentence rule).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* codex_roles.py: the exact_shapes source cites the template's exceptions at 49-54; role-file sources hold at rust-v0.159.2

Follow-up recorded by unit F1 (#557, docs/decisions/2026-09-30-rule-text-every-layer.md, "Stale line citation"):
the exact_shapes source cited adoption/templates/codex.AGENTS.template.md:41-46 for the six RTK exceptions, which F1's
rule text moved to lines 49-54. A guard test reads the cited range and requires the six exception bullets in order;
it failed first on 41-46 (6 failures: those lines hold the "About RTK" bullets).

The upstream citation at the branch's codex_roles.py:357, codex-rs/agent-roles/src/agent_role_config.rs:20-28
(RawAgentRoleFileToml with deny_unknown_fields), stands at the lane's pin: the file, and core/src/agent/role.rs, are
byte-identical at openai/codex rust-v0.157.1 and rust-v0.159.2 (sha256 70ba8cf41c7339a0... and 0311e6438eda278a...,
read 2026-10-01 from both tags' raw files). The RULES comment and the F4 addendum's Sources record that.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* MCP template check reads every SubagentStart carrier block; the template registers exactly their servers

Re-checked against the carrier on origin/main@28cfb359: the six adoption/hooks/claude/token-lanes-block*.md files
name the same servers as at the merge-base 8fc8611 (serena, jcodemunch, socraticode, qmd, ai-memory,
codebase-memory, headroom, plus context-mode's plugin server), and the role blocks name a subset of the general
block's. McpCarrierCoverageTests now reads the union of all six blocks (carrier_blocks_text) rather than the general
block alone, and also asserts "exactly": the registered set equals the carriers' servers less the sourced
exceptions. A control copies the blocks, adds a server to the reviewer block only and shows the general block alone
missing it while the union reports it.

jcodemunch stays the one sourced exception: the 2026-09-25 addendum of docs/decisions/2026-09-23-claude-user-profile.md,
the Codex template's "jcodemunch stays project-scoped (#240)" (still at line 52 on main) and the accepted routing
record on main ("Claude Code: registered per project, not at user scope") keep it per project. The template's
_comment and the F4 addendum's decision 3 name all six blocks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* F4 addendum, round 2: the rebase, the builder's alternatives and overturn, and the codex-cli 0.159.2 dry run

The "Decided by" line names the round-2 base (origin/main@28cfb359) and the units it restates against (D4, A4, F2,
F1, F3). Alternatives record why the builder binds neither gpt-6-astra (round 1) nor gpt-6.1-sol, and why
${CODEX_MODEL} cannot stand in for a role file. The overturn condition says when the builder takes a model again.

Evidence, local integration at the lane's pin: the pinned codex-cli 0.159.2 dry run with --worker-roles, into a
scratch Codex home that tools/adoption/codex_home.py made from the rendered user template (adoption/hosts/example.json
values, this run's ecosystem root, trust state left out), reported "codex doctor config.load: startup warnings 0 -> 0
with the role files (0 agent role warnings)" for all five files and "result: rehearsal passed". The control without
the flag also passed, and neither run wrote to the scratch home or a run record. Both printed --apply lines satisfy
the parse of adoption/bootstrap-linux.sh:1000-1001. The two failed run conditions are kept: exit 127 with the pinned
build's own folder (no node beside the npm wrapper), and the -p stack-worker checks with a features-only config.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* codex_roles.py: cite the template's six RTK exceptions by their marker, after #568 moved them again

Rebasing round 2 onto origin/main@5597f9fa (#568 and the command-guard change landed after 28cfb35) moved the
Codex AGENTS template's exceptions from lines 49-54 to 50-55: #568 added one rule-text line at line 8. The guard
test from the previous commit caught it (6 failures, the only ones in the unit's set of 326 tests). Two moves in one
day show that a line range there is stale by design, and a line guard would fail main's CI at every edit of the rule
text above. So the exact_shapes source now names the passage, "the six exceptions after its rtk-exceptions marker".
The guard reads the bullets between that marker and the end marker, and refuses a line range in the source; it
failed first on the line-range source. The F4 addendum's round-2 line names the new base and the move.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* F4 addendum: the alternatives and the Gate A flip list cite the lines of round 2's head

Round 1 cited its base's lines. Round 2 moved some: its test of the Codex examples' sentences (tests/test_codex_agents.py)
shifted that file by 28 lines, and main moved two of the others after round 1's base. Restated and checked line by
line at this head: tests/test_codex_agents.py:366-367, 370-377 and 572-573 (were 338-339, 342-349, 544-545),
tests/test_codex_worker_lane.py:144 and 1043 (were 140 and 1001), scripts/adoption_status.py:224 (was 194).
tools/adoption/prove_codex_lane.py:149-173, tools/token-e2e/freeze_snapshot.py:108 and :1244 and the examples
README's lines still hold. Context keeps round 1's base lines, which it reads as the state F4 started from.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Routing record: its line convention names the revision F4's restated rows read

The rows "GPT-6 judgment roles" and "Generic Codex children" now cite worker-role lines "as read at" #548's head,
which the Decision's statement of where line numbers are read did not name. Text only; the record is not hash-listed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* F4 round 3: the jcodemunch exception requires Claude Code's per-project registration; no-flag sentences corrected

Cross-family review of round 2 (cx/gpt-6.1-sol, max, whole branch at 112bd68): needs_changes, one medium, one low.

- tests/test_install_claude_profile.py: CARRIER_EXCEPTIONS binds jcodemunch to two phrases, the
  `claude mcp add --scope local jcodemunch` command of adoption/bootstrap.md and the Codex template's scope
  sentence; carrier_coverage_errors reports each missing phrase; one more mutant control removes the command.
  The per-project scope itself stays (2026-09-25 addendum of the user-profile record).
- docs/decisions/2026-09-26-stack-agents-role-dispatch.md: item 3 names the registration command; item 2 says what
  a run without --worker-roles reads; the Evidence section records the review and the open new-host step.
- tools/adoption/apply_codex_lane.py: the comment at the worker-role pins says the same.

Tests: python3 -m unittest tests.test_install_claude_profile tests.test_codex_roles tests.test_codex_agents
tests.test_codex_worker_lane tests.test_adoption_docs_consistency tests.test_task_model_routing -> 291 tests OK
(15 skipped), exit 0.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants