Repository navigation
Token-efficiency profile accepted with the three code-navigation tools; task-to-model routing record (unit A4) - #540
Conversation
… code-navigation tools stay optional rows Round 1 (24c8703) added jcodemunch-mcp, codebase-memory-mcp and ast-grep to the profile's component_ids and required_commands. Neither adoption/pins-linux-x86_64.json nor adoption/pins-macos-arm64.json lists them, and both bootstraps fail closed on a selected component with no pin: adoption/bootstrap-linux.sh:169 and adoption/bootstrap-macos.sh:220 print "No pin in <file> for selected component(s)" and exit 3, and adoption/bootstrap-macos.sh:176-190 records that no component is exempted from a pin by default. PR #540's validate-macos job failed on it (3 tests of tests/test_adoption_bootstrap_macos.py TokenEfficiencyPlanTests), reproduced locally at 8ca7895 with `python3 -m unittest tests.test_adoption_bootstrap_macos tests.test_adoption_bootstrap tests.test_adoption_launchd`: FAILED (failures=3, skipped=38). The profile is again the 14 pinned components. It keeps the accepted label and the routing record in recipe_paths. The three tools stay the code-navigation layer's task-selected tools and the carrier's task-appended lanes, installed on demand from their recipes, until each has a reviewed pin on both platforms and a host receipt. tests/test_adoption_status.py returns to main's OPTIONAL set of eight and the (16, 14, 10) Ultracode split, and TokenEfficiencyProfileTests gains test_every_profile_component_has_a_pin_on_both_platforms, so a profile that outruns the pin files fails in this file's own module set and not only in the bootstrap plan tests. tests/test_task_model_routing.py holds the record's claims about the three: outside component_ids and required_commands, unpinned on both platforms, named by the code-navigation current choice and by the carrier's lane ids. Failing-first: with these tests and the round-1 manifest (17 component_ids), 4 of the 12 tests in TokenEfficiencyProfileTests and TaskModelRoutingRecordTests fail, 8 failures counting subtests ("Tuples differ: (16, 17, 13) != (16, 14, 10)", "'jcodemunch-mcp' unexpectedly found in [...]", the pin-file membership checks). With the manifest of this commit all 12 pass. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
8ca7895 to
c8f934c
Compare
|
Round 2 (head replaced under the hot-file protocol; pushed with --force-with-lease): Reverting the round-1 profile change. CI fact: run 36728291629 (head 8ca7895), job validate-macos, step "Gate on the adoption test modules" failed: Both bootstraps fail closed on a selected component with no pin ( |
|
Cross-family review of the repaired head c8f934c (GPT-6 through the OmniRoute gateway, read-only, diff against the merge-base 11227bf): Coordinator adjudication: the finding restates the unit brief's original deliverable 1 (add the three code-navigation tools to the profile). That deliverable is superseded on evidence: both bootstraps fail closed on any selected component without a pin ( |
c8f934c to
d004de1
Compare
… code-navigation tools stay optional rows Round 1 (24c8703) added jcodemunch-mcp, codebase-memory-mcp and ast-grep to the profile's component_ids and required_commands. Neither adoption/pins-linux-x86_64.json nor adoption/pins-macos-arm64.json lists them, and both bootstraps fail closed on a selected component with no pin: adoption/bootstrap-linux.sh:169 and adoption/bootstrap-macos.sh:220 print "No pin in <file> for selected component(s)" and exit 3, and adoption/bootstrap-macos.sh:176-190 records that no component is exempted from a pin by default. PR #540's validate-macos job failed on it (3 tests of tests/test_adoption_bootstrap_macos.py TokenEfficiencyPlanTests), reproduced locally at 8ca7895 with `python3 -m unittest tests.test_adoption_bootstrap_macos tests.test_adoption_bootstrap tests.test_adoption_launchd`: FAILED (failures=3, skipped=38). The profile is again the 14 pinned components. It keeps the accepted label and the routing record in recipe_paths. The three tools stay the code-navigation layer's task-selected tools and the carrier's task-appended lanes, installed on demand from their recipes, until each has a reviewed pin on both platforms and a host receipt. tests/test_adoption_status.py returns to main's OPTIONAL set of eight and the (16, 14, 10) Ultracode split, and TokenEfficiencyProfileTests gains test_every_profile_component_has_a_pin_on_both_platforms, so a profile that outruns the pin files fails in this file's own module set and not only in the bootstrap plan tests. tests/test_task_model_routing.py holds the record's claims about the three: outside component_ids and required_commands, unpinned on both platforms, named by the code-navigation current choice and by the carrier's lane ids. Failing-first: with these tests and the round-1 manifest (17 component_ids), 4 of the 12 tests in TokenEfficiencyProfileTests and TaskModelRoutingRecordTests fail, 8 failures counting subtests ("Tuples differ: (16, 17, 13) != (16, 14, 10)", "'jcodemunch-mcp' unexpectedly found in [...]", the pin-file membership checks). With the manifest of this commit all 12 pass. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
Rebased onto 🤖 Generated with Claude Code |
… code-navigation tools stay optional rows Round 1 (24c8703) added jcodemunch-mcp, codebase-memory-mcp and ast-grep to the profile's component_ids and required_commands. Neither adoption/pins-linux-x86_64.json nor adoption/pins-macos-arm64.json lists them, and both bootstraps fail closed on a selected component with no pin: adoption/bootstrap-linux.sh:169 and adoption/bootstrap-macos.sh:220 print "No pin in <file> for selected component(s)" and exit 3, and adoption/bootstrap-macos.sh:176-190 records that no component is exempted from a pin by default. PR #540's validate-macos job failed on it (3 tests of tests/test_adoption_bootstrap_macos.py TokenEfficiencyPlanTests), reproduced locally at 8ca7895 with `python3 -m unittest tests.test_adoption_bootstrap_macos tests.test_adoption_bootstrap tests.test_adoption_launchd`: FAILED (failures=3, skipped=38). The profile is again the 14 pinned components. It keeps the accepted label and the routing record in recipe_paths. The three tools stay the code-navigation layer's task-selected tools and the carrier's task-appended lanes, installed on demand from their recipes, until each has a reviewed pin on both platforms and a host receipt. tests/test_adoption_status.py returns to main's OPTIONAL set of eight and the (16, 14, 10) Ultracode split, and TokenEfficiencyProfileTests gains test_every_profile_component_has_a_pin_on_both_platforms, so a profile that outruns the pin files fails in this file's own module set and not only in the bootstrap plan tests. tests/test_task_model_routing.py holds the record's claims about the three: outside component_ids and required_commands, unpinned on both platforms, named by the code-navigation current choice and by the carrier's lane ids. Failing-first: with these tests and the round-1 manifest (17 component_ids), 4 of the 12 tests in TokenEfficiencyProfileTests and TaskModelRoutingRecordTests fail, 8 failures counting subtests ("Tuples differ: (16, 17, 13) != (16, 14, 10)", "'jcodemunch-mcp' unexpectedly found in [...]", the pin-file membership checks). With the manifest of this commit all 12 pass. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
d004de1 to
69f3e9e
Compare
|
Stacked merge train after #542 merged (main 1f2cdce): this branch is rebased onto #545 3805a59 with the hot-file protocol applied against that head (its evidence.json taken, this unit's files re-registered in the last commit; 🤖 Generated with Claude Code |
|
Accepted foundation runtime enhancement handoff from PR551, head The Claude dispatcher now defaults to the enhanced scoped native SDK home and gates model execution on readiness. Actual bounded acceptance completed: Sol/Max read the selected skill, Context Mode counted the unchanged test, Serena returned Original-field receipt, scoped record, and native kit. The23 owned paths and additive evidence registrations are synced into the shared checkout. Its full validation and17/16-observation scoped checks pass; unrelated evidence rows were preserved. Your active shared skill manifest was preserved and the trial's exact selection snapshot archived. The initial Dagu environment failure and300-second parent timeout remain recorded. Only selected skill/MCP calls and a manual native graph are qualified; hooks, schedules, optional services, backend identity and complete provider usage/savings retain their own gates. Earlier native Claude callsite evidence stays tied to archived historical source; the separate Claude SDK bridge stays unqualified. Token/checkpoint handoff: successful native cumulative scopes are Sol parent153,813 and Astra child66,044, each counted once. Cached input/reasoning are subsets; observer completion sums corroborate those counters and must not be added again. All nine successful-run unclassified flow_error records are retained. Failed-parent usage remains a partial native snapshot, and whole-task savings remains |
Round 2: routing record reconciled with D4 (#542), restacked on F1Head Summary. The record and its test predated D4 (merged as
SOTA sources
EvidenceAll checks are our integration checks; no upstream test and no model run is involved.
Known residuals: the record cites the CI run by the superseded head 🤖 Generated with Claude Code |
… code-navigation tools stay optional rows Round 1 (24c8703) added jcodemunch-mcp, codebase-memory-mcp and ast-grep to the profile's component_ids and required_commands. Neither adoption/pins-linux-x86_64.json nor adoption/pins-macos-arm64.json lists them, and both bootstraps fail closed on a selected component with no pin: adoption/bootstrap-linux.sh:169 and adoption/bootstrap-macos.sh:220 print "No pin in <file> for selected component(s)" and exit 3, and adoption/bootstrap-macos.sh:176-190 records that no component is exempted from a pin by default. PR #540's validate-macos job failed on it (3 tests of tests/test_adoption_bootstrap_macos.py TokenEfficiencyPlanTests), reproduced locally at 8ca7895 with `python3 -m unittest tests.test_adoption_bootstrap_macos tests.test_adoption_bootstrap tests.test_adoption_launchd`: FAILED (failures=3, skipped=38). The profile is again the 14 pinned components. It keeps the accepted label and the routing record in recipe_paths. The three tools stay the code-navigation layer's task-selected tools and the carrier's task-appended lanes, installed on demand from their recipes, until each has a reviewed pin on both platforms and a host receipt. tests/test_adoption_status.py returns to main's OPTIONAL set of eight and the (16, 14, 10) Ultracode split, and TokenEfficiencyProfileTests gains test_every_profile_component_has_a_pin_on_both_platforms, so a profile that outruns the pin files fails in this file's own module set and not only in the bootstrap plan tests. tests/test_task_model_routing.py holds the record's claims about the three: outside component_ids and required_commands, unpinned on both platforms, named by the code-navigation current choice and by the carrier's lane ids. Failing-first: with these tests and the round-1 manifest (17 component_ids), 4 of the 12 tests in TokenEfficiencyProfileTests and TaskModelRoutingRecordTests fail, 8 failures counting subtests ("Tuples differ: (16, 17, 13) != (16, 14, 10)", "'jcodemunch-mcp' unexpectedly found in [...]", the pin-file membership checks). With the manifest of this commit all 12 pass. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
69f3e9e to
901ed76
Compare
|
Cross-family review of the round-2 delta ( |
Round 3: GPT-6.1 routing checked binding by binding (GPT-6 review of 901ed76: needs_changes, medium)Head Summary. The "nowhere else" check compared files only, so a
SOTA sources
EvidenceAll checks are our integration checks; no upstream test and no model run is involved.
Residuals: the test now carries a second copy of D4's allowed bindings ( 🤖 Generated with Claude Code |
901ed76 to
e3acc20
Compare
…nd token lanes Adds the 2026-09-30 standing clauses to the top-rule block of adoption/templates/codex.AGENTS.template.md: skill discovery (search-first, find-skills, skill-creator; implicit invocation from a skill's description and explicit $skill-name, per openai/codex rust-v0.159.2 codex-rs/ext/skills/src/catalog_prompt.rs:8), upstream A/B and E2E harnesses, the completeness critic feeding the next landscape sweep, the north-star direction, GPT-6 model routing through the OmniRoute gateway (codex -p omniroute), dated decision records, the startup rule, and a token-lanes line naming each MCP server the Codex config template registers. The RTK upstream text and the exceptions block stay byte-identical, so the F4 block in every Codex role is unchanged. The top-rule pin is re-derived with template_segments(): 455 words, 7b41478f...; a new test requires a lane for every server in codex.config.template.toml and the standing phrases. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tanding clauses examples/claude-native/CLAUDE.md becomes the single managed source of the operator's user-level file. Step 1 of the top rule takes the user-level paragraph (research and record, reference implementations, popularity guides discovery) and the skill-discovery clause (search-first, find-skills with `npx skills find`, skill-creator); the core rule takes the rules only the user-level file held (short plan, pins and reasons, relevant layers, reviewer agreement is not proof), the north-star direction, the upstream A/B and E2E harnesses and the completeness critic; token practice takes the startup rule; the worker section takes GPT-6 routing through the OmniRoute gateway. Every worker, model, Ultracode and agent-team rule is kept word for word. PortableTopRuleTests gains a phrase check for the clauses and the user-level rules (it failed on the unedited template with 29 phrases missing, 8 of them rules the user-level file held) and re-baselines the word budget to 1,656. @RTK.md stays the host's own import (recipes/claude-native-profile.md:134). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…gence log row docs/decisions/2026-09-30-rule-text-every-layer.md supersedes the 2026-09-28 top-rule record's "the operator's user-level file is the operator's own" at the user's 2026-09-30 request: the portable template becomes the single managed source of that file. It maps the six clauses to each surface, records the alternatives (a UserPromptSubmit carrier, per the hooks reference, adds its context beside every prompt; leaving the file unmanaged produced the divergence; per-agent copies are unit F2's; pointers measured as the fallback), the o200k sizes from token_manifest.count_files, the overrun of the unit's 200-token budget with its floor and fallback, and the overturn conditions. docs/harness-defaults.md logs the host/template divergence, naming only the check seen failing on it (the new PortableTopRuleTests phrase check). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ter hashes Hot files last (docs/lanes.md hot-file protocol). AGENTS.md takes the standing clauses by amending, not appending, the lines the brief names: - line 3, the top rule: skill discovery (search-first, find-skills with `npx skills find`, skill-creator; manifest skills model-invocable in both clients) replaces "with the installed research and skill-discovery skills", and A/B and E2E on upstream harnesses, never a self-written runner; - line 7: the north-star R&D direction; each unit names its action; - line 18: the completeness critic feeding the layer's next landscape sweep; - line 28: GPT-6 through the OmniRoute gateway (astra max, sol medium, 6.1 sol pending qualification), Codex CLI as the second native client, and "no audits, trials or network at startup" with the daily currency timer's due-file line (record lands with unit A2) replacing "Do not rerun the full audit or model trials at startup". manifests/evidence.json re-registers the six changed files already listed, with scripts/host_receipts.py register_file; component_matrix and new_host_grand_list --check pass unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The phrase check's committed form (with "from the selected source revision", the user-level file's own wording) finds 29 phrases missing from the unedited template, 9 of them held only by the user-level file; the record, its Checks section and the anti-pattern row said 8 (from a run before that phrase changed). The record also states the two non-verbatim worker lines exactly: "Quality comes first" is one line with an identical word sequence, and the sizing line keeps every user-level word and adds "to its task". Re-registers docs/harness-defaults.md (hot file, last commit). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
"No automatic router exists" now also names Claude Code's own content-based fallback and the two keys that switch it off in the project settings and the user-settings template (switchModelsOnFlag false, CLAUDE_CODE_DISABLE_REFUSAL_FALLBACK=1), with the model-currency addendum's note that their coverage is unverified and that availability fallback chains are a separate mechanism. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#508's token-stack record says the run-shape levers, role dispatch among them, are "owned by the Gate A owner after #381 closes" (line 90). Its line 93 ("Gate A per-row results decide any change") is about the on-demand members of line 92, not the levers, so the Context paragraph and the Sources entry now cite line 90 only (review finding, low). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…odex rows name no version Six sibling units of this wave may edit files that the record quotes by line: landscape-sweep (A1), bootstrap-linux.sh (A3), .claude/agents and the settings template's hooks entry (F2), the settings template and the Codex template (F3), codex_roles.py (F4) and the model-currency record (D4, an addendum over the region of its line 284). None of them may edit this record or its test. An exact-line check fails their pull requests on any inserted line, even when no route changes. - The record now states every line number as read at e45328d. tests/test_task_model_routing.py counts each quoted value in its whole file and requires at least one copy for each distinct line quoted, so a pure line shift passes and a changed model or effort fails. - Two values have more copies than the record cited. It now also cites the template's modelSettings xhigh for claude-opus-5-5 and claude-sonnet-5-5 (:303, :306) and review-changes.js's recheck-stage scout binding (:78), so a change to any cited copy fails. - model-currency.md:284 is cited without a CI-checked quote. D4 may rewrite that sentence. - The Codex rows' Client cell says "Codex CLI". The preamble records that manifests/stack.json:297 pins 0.157.1 at e45328d, so D4's pin move leaves no stale version string (review finding, low). - The GPT-6.1 check matches a model binding (a model key or constant, or -m/--model with an optional gateway prefix), not any mention. F1 may write prose naming gpt-6.1-sol into adoption/templates; a real binding, such as D4 switching the Codex template default, still fails. Controls, each run in a fresh copy of the files the test reads: 16 of 16 as expected for both the old and the new test. Pure line shifts (codex template, model-currency.md, sweep.js) and an F1-style prose mention fail the old test and pass the new one. Changed frontmatter, one of two identical sweep bindings, and a gpt-6.1-sol template default fail both. Changing the :306 per-model xhigh or the :78 recheck binding passes the old test and fails the new one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…navigation tools The profile's label drops "Drafted, not accepted" for a dated acceptance of the selection, citing the routing record, which joins recipe_paths so the profile resolves only while the record is present. jcodemunch-mcp, codebase-memory-mcp and ast-grep join component_ids and required_commands: the SubagentStart carrier names them as task-appended lanes and the code-navigation layer's current choice names all three (catalogs/landscape/foundation.json). Their versions stay in manifests/stack.json (1.108.319, 0.11.0, 0.45.3); a manifest profile has five keys and carries no version or wiring field, so neither is added. tests/test_adoption_status.py restates the split the profile is checked against: NAVIGATION_CHOICE, asserted against the code-navigation current choice, leaves OPTIONAL, and the Ultracode tool split becomes 13 profile rows and 3 optional rows against 17 component_ids. Failing-first: the restated tests fail against the previous manifest (2 of 3 in TokenEfficiencyProfileTests) and the previous tests fail against this manifest (the same 2, plus ProfileTableTests.test_pin_columns_match_the_pin_files, whose adoption/README.md cells are outside this unit's paths). manifests/evidence.json is re-registered in the branch's last commit. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…tools' wiring The record no longer says the profile is "not decided here"; it makes the claim that adoption/manifest.json now cites. The Decision gains a dated acceptance that bounds itself to the selection and its routing (structural validation, not a host's acceptance), the pin and the Claude Code and Codex wiring of jcodemunch-mcp, codebase-memory-mcp and ast-grep with the file that carries each, the basis for adding them (the SubagentStart carrier, the code-navigation current choice, the 2026-09-25 Ultracode run), and where #508 differs: it lists ast-grep as the structural-code lane, makes jCodeMunch an owner only once wired and lists codebase-memory-mcp as not a member, so the record does not rest on #508 for those two. Two alternatives (leave the three optional; a new manifest key) and two overturn conditions (#508 or Gate A changes a row; pins arrive) are added, and the Sources name the new files and #508 lines 85, 86 and 103. Every path:line cite in the whole record resolves in this branch (115 cites, 28 bare paths); the record keeps its five sections and one table, which tests/test_task_model_routing.py holds. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…tools now in the profile Three statements had become false: that jCodeMunch, ast-grep and codebase-memory-mcp are optional rows outside the profile, that the profile has 14 component_ids, and that the Ultracode run's 16 tools are ten profile rows and six optional rows. The profile paragraph now lists the three as the code-navigation layer's task-selected tools (none a layer winner), the optional-row paragraph keeps Context Hub, the viewers and OmniRoute, and adds why client_wiring still checks only Serena, SocratiCode and ai-memory, that the three have no platform pin (so --pinned-versions reports them unchecked) and that a bootstrap refuses the profile until they are named in --allow-unpinned (adoption/bootstrap-linux.sh --help: exit 3). The Ultracode paragraph becomes 17 component_ids, thirteen profile rows and three optional rows, matching tests/test_adoption_status.py. The review of this unit asked for this reconciliation (docs/token-efficiency-stack.md lines 264-270); the file is outside the brief's allowed paths, so it is its own commit. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…epts it The record's Decision says the profile's label accepts it and names the record, that the record is one of its recipe_paths, and that jcodemunch-mcp, codebase-memory-mcp and ast-grep are component_ids and required_commands. The new test in tests/test_task_model_routing.py holds those claims to adoption/manifest.json and requires the record to name each tool. It compares no version: manifests/stack.json owns them, and the record now dates the ones it quotes at e45328d. Failing-first: against the previous manifest the test fails on the label ('Accepted' not found in 'Drafted, not accepted: ...'); against this branch's manifest all 8 tests in the module pass. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…d re-registration (companion) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… code-navigation tools stay optional rows Round 1 (24c8703) added jcodemunch-mcp, codebase-memory-mcp and ast-grep to the profile's component_ids and required_commands. Neither adoption/pins-linux-x86_64.json nor adoption/pins-macos-arm64.json lists them, and both bootstraps fail closed on a selected component with no pin: adoption/bootstrap-linux.sh:169 and adoption/bootstrap-macos.sh:220 print "No pin in <file> for selected component(s)" and exit 3, and adoption/bootstrap-macos.sh:176-190 records that no component is exempted from a pin by default. PR #540's validate-macos job failed on it (3 tests of tests/test_adoption_bootstrap_macos.py TokenEfficiencyPlanTests), reproduced locally at 8ca7895 with `python3 -m unittest tests.test_adoption_bootstrap_macos tests.test_adoption_bootstrap tests.test_adoption_launchd`: FAILED (failures=3, skipped=38). The profile is again the 14 pinned components. It keeps the accepted label and the routing record in recipe_paths. The three tools stay the code-navigation layer's task-selected tools and the carrier's task-appended lanes, installed on demand from their recipes, until each has a reviewed pin on both platforms and a host receipt. tests/test_adoption_status.py returns to main's OPTIONAL set of eight and the (16, 14, 10) Ultracode split, and TokenEfficiencyProfileTests gains test_every_profile_component_has_a_pin_on_both_platforms, so a profile that outruns the pin files fails in this file's own module set and not only in the bootstrap plan tests. tests/test_task_model_routing.py holds the record's claims about the three: outside component_ids and required_commands, unpinned on both platforms, named by the code-navigation current choice and by the carrier's lane ids. Failing-first: with these tests and the round-1 manifest (17 component_ids), 4 of the 12 tests in TokenEfficiencyProfileTests and TaskModelRoutingRecordTests fail, 8 failures counting subtests ("Tuples differ: (16, 17, 13) != (16, 14, 10)", "'jcodemunch-mcp' unexpectedly found in [...]", the pin-file membership checks). With the manifest of this commit all 12 pass. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ee tools are not profile members This reverts 12a587e. docs/token-efficiency-stack.md is byte-identical to origin/main (11227bf) again: the profile has 14 component_ids, jCodeMunch, ast-grep and codebase-memory-mcp are task-selected options in the code-navigation layer's current choice and stay optional rows, and the Ultracode run's 16 tools are ten profile rows and six optional rows. The statements 12a587e rewrote (three joined the profile on 2026-09-30; a bootstrap refuses the profile until they are named in --allow-unpinned; 17 component_ids, thirteen profile rows) described a profile that adoption/manifest.json no longer has, because both bootstraps fail closed on a component without a pin and none of the three has one. tests/test_adoption_status.py checks the ten/six split against the receipt's tool list. The reasoning is recorded once, in docs/decisions/2026-09-30-task-model-routing.md. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…s describe the 14 pinned components again d5369cb wrote the row for 17 members: "14 of 17 (all 14 at v2026.09.26.2)" in both pin columns, the three code-navigation tools in the description, and a note that a bootstrap of the profile needs them in --allow-unpinned. The profile has its 14 pinned components again (previous commit), so the description, both pin cells and the note are main's, and only the acceptance wording stays: "Accepted 2026-09-30 as the selection", linking the routing record. The pinned release v2026.09.26.2 covers the same 14 of 14 as HEAD, so tests/test_adoption_docs_consistency.py (ProfileTableTests.cell_errors) asks for no note on that tag; the notes for v2026.09.25.2 and v2026.09.26 are checked against their own tags and hold. Failing-first: with the round-1 row and the 14-component manifest, ProfileTableTests.test_pin_columns_match_the_pin_files fails ("token-efficiency Linux pins: table says '14 of 17', adoption/pins-linux-x86_64.json gives 'all 14'"). With this row the module passes: Ran 37 tests, OK (skipped=1). The one skip is data dependent, not environmental: no profile's coverage differs between the pinned release and HEAD, on origin/main as well as here. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…pted profile, for want of a pin The record accepted the token-efficiency profile and, in round 1, added jCodeMunch, codebase-memory-mcp and ast-grep to it. It now accepts the profile as it is, its 14 pinned components, and says why the three are not members: both bootstraps fail closed on a selected component with no pin (adoption/bootstrap-linux.sh:144-172, adoption/bootstrap-macos.sh:177-223, and the macOS comment at :177-187 that no component is exempted from a pin by default any more), none of the three has an entry in adoption/pins-linux-x86_64.json or adoption/pins-macos-arm64.json, and tests/test_adoption_bootstrap_macos.py:1015-1017 requires the profile's plan to resolve every component with no --allow-unpinned. - Alternatives: leaving the three out is the chosen state; adding them (round 1) is a rejected alternative with its failure (Adoption bootstrap smoke run 36728291629, job validate-macos, step "Gate on the adoption test modules"; reproduced locally at 8ca7895, failures=3) and the reason --allow-unpinned does not rescue it. - Decision: "Three tools join the profile" becomes "Three tools stay outside the profile"; the wiring bullets, which describe how each installs on demand, stay; the "14 of 17" coverage and the exit-3 sentence are gone; "Where #508 differs" becomes "#508's rows": leaving the three out agrees with #508 for jCodeMunch (line 86) and codebase-memory-mcp (line 103), and ast-grep (line 85, the structural-code lane) is out only for want of a pin. - Overturn condition: the follow-up unit that adds a tool once it has a reviewed pin on both platforms and a host receipt for each, with the files it must restate in the same change (adoption/README.md row, token-efficiency-stack.md split, OPTIONAL and CARRIER_TOOLS in the two tests). - Header and Sources: the citations are at e45328d while the branch is based on 11227bf, where examples/claude-native/workflows/README.md has grown by 249 lines; the record says so. The pin-rule files it now cites are unchanged between the two bases. tests/test_task_model_routing.py (previous commit) holds these claims to the files. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
tests/test_task_model_routing.py predated D4 (#542, merged as 1f2cdce) and failed on the stacked tree (validate job 110072600341): two quoted `model = "gpt-6-astra"` lines and the test that expected GPT-6.1 Sol to be routed nowhere. - A quoted value counts only as a whole value, so `model = "${CODEX_MODEL}"` is not met by the template's `default_subagent_model` line. - GPT-6.1 Sol: the table's Sol rows are exactly the Sol-primary record's three routes (the Codex coordinator at ultra, primary workers and generic children at max), and every file in the routing places (now with recipes/ and adoption/agents/) that binds a GPT-6.1 model, a constant such as CODEX_MODEL_CURRENT included, is cited by a Sol row. __pycache__ and node_modules are skipped: CI flagged prove_codex_lane's bytecode. - The Codex user template is rendered through tools/adoption/render_config.py for every pinned platform, as tests/test_render_config.py's CodexModelTests do, and must give the models and efforts that the coordinator and generic-children rows name. - The task classes add complex-workflow coordination and a single consequential judgment. Fails against the record at this commit (5 failures, 1 error); the next commit restates it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
D4 made GPT-6.1 Sol the Codex coordinator (ultra) and primary-worker model (max), with GPT-6 Astra for complex-workflow coordination (ultra) and a single consequential judgment (max): docs/decisions/2026-09-30-sol-primary-quality-defaults.md and the model-currency addendum of 2026-09-30. The record still listed Sol as pending and routed nowhere. - Six Codex rows, read at 1f2cdce: the Codex judgment role carriers (Astra, max); the coordinator and interactive default (the template's ${CODEX_MODEL} at ultra, rendered as gpt-6.1-sol for the Linux pin 0.159.2 and gpt-6-astra for the macOS pin 0.155.1); primary workers (the stack-worker profile, the lane's worker pins and live proof, the recipe's command); generic children; complex-workflow coordination and escalation (instruction only). - Context and Decision state the D4 routing and its dependence on the Codex pin. No automatic router exists, and D04/D06/D11 stay off until the D3 promptfoo A/B, as before. - Overturn: the user's next decision or D4's reopening condition, and a Codex pin crossing 0.159.1 on a platform. Sources: D4's records, the enforcement points and the CI job. The Claude rows are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mand profile The GPT-6 cross-family review of round 2 (901ed76, needs_changes, medium): the "nowhere else" check compared files only, so a `[profiles.security-review]` table with `model = "gpt-6.1-sol"` added to the user template, a file the coordinator row already cites, left all nine routing tests passing. - Bindings are counted one by one: a TOML key by its table (parsed with tomllib: `model`, `agents.default_subagent_model`, `profiles.<name>.model`, on the ID or on `${CODEX_MODEL}`), a key or constant elsewhere (CODEX_MODEL_CURRENT), and a command's `-m`, named with the profile the command selects (`-p stack-worker -m`). - Each binding must be one a Sol row cites by file and name. A quoted TOML key names the one key of that name and value (a quoted `[table]` header scopes it), and a quote that fits two keys is reported. Each cited binding must be one D4 makes: the template's `model` and `[agents] default_subagent_model`, the render constant, the stack-worker profile's `model` and the stack-worker command's `-m`. - Regression: the planted profile, in memory, on the ID and on the placeholder, and then cited by its table from the coordinator row; each variant is reported. Fails against the record at this commit: the stack-worker command's `-m` in the profile's header (line 3) and in the live proof's help (tools/adoption/prove_codex_lane.py:36) are bindings no Sol row cites. The next commit cites them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r and the live proof write it The binding-level check of the previous commit found two GPT-6.1 bindings that the primary-worker row did not cite, both the stack-worker command's `-m`, which D4 routes: the profile's header (adoption/templates/codex.stack-worker.config.toml:3) and the live proof's help (tools/adoption/prove_codex_lane.py:36). The row now quotes both, beside the recipe's command it already quoted. No other row changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…(hot-file protocol) The last commit of the branch carries the hot files: the stacking ref's manifests/evidence.json with this branch's changed hash-listed files re-registered, and the regenerated component-evidence matrix and new-host grand list (docs/lanes.md hot-file protocol). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
e3acc20 to
df92df6
Compare
|
Concrete foundation runtime follow-up is published as PR566, head The19-role packet resolves the broader landscape from352 public stars/four maintained awesome lists into task-specific choices. Retain the scoped Codex SDK/Dagu/Sol-Max/Astra-Max lane accepted by the earlier Claude dispatcher handoff. New native OpenHands CLI1.16.0, SDK/tools1.50.1 and DeepAgents0.7.21 recipes are separate from your frozen O1/global default work. Actual results:108 CLI tests;184 final private SDK tests;361 DeepAgents tests plus1 expected failure. One native OpenHands Sol-Max Responses exact-output request passed. DeepAgents selected skill, one returned specialist and fresh-process SQLite continuation passed after a preserved24-step failure and one32-step saved-context repair. Same-oracle negative controls failed as intended. Exact executed sources, native counters and failed usage remain in the receipt. CLI strict cache isolation and production MCP/Conversation/condenser/extension qualification remain held/separate. Primary review paths:
Completed Astra architecture and independent OH setup reviews remain scoped. Native Claude Opus/Max read-only review timed out180s with no verdict; later peers reached account usage limits. Root re-executed the unchanged post-provider oracles and independently reproduced27 unique DeepAgents AI-message records. No completed broader peer acknowledgement is claimed. For the dashboard owner, the concrete checkpoint payload is:
Please reconcile these rows through your owned shared checkpoint and review the concrete role/default boundaries when the native peer is available. Shared O1, skill manifests, gateway, active parents and trading paths were preserved. |
…routing record lists the worker roles Stacked-state check of #548 against unit D4 (#542) and unit A4's routing record (#540) on origin/main@28cfb359. tests/test_task_model_routing.py passes: no worker role, the applier or codex_roles.py binds GPT-6.1 or ${CODEX_MODEL}. The builder's gpt-6-astra binding is the gap that test does not scan. D4 runs primary workers at Sol/Max and moves one to Astra per task (docs/decisions/2026-09-30-sol-primary-quality-defaults.md:13-20,27-30) and preserves Astra for judgment roles (:21-22). openai/codex rust-v0.159.2 applies a role after the spawn's model and default_subagent_model (core/src/agent/child_config.rs:62-73,204-206; core/src/agent/role.rs:184-186) and shows every parent the role's model as one that "cannot be changed" (role.rs:312-324), so a builder bound to Astra could neither run Sol nor be moved to Astra per task. isolated-builder.toml names no model and keeps max. codex_roles.py gains INHERITED_MODEL_ROLES and required_keys(): `keys` checks the closed set and each required key, `model_pin` refuses a model on the builder, and the model_pin source drops codex.stack-worker.config.toml:12, which D4 made gpt-6.1-sol. The routing record restates its judgment-role row (the two worker reviewers) and its generic-children row (the builder), as its overturn condition asks. Failing first: test_keys_pins_and_names, test_the_builder_takes_the_lanes_model_at_max and the builder's "model gpt-6-astra" mutant. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…routing record lists the worker roles Stacked-state check of #548 against unit D4 (#542) and unit A4's routing record (#540) on origin/main@28cfb359. tests/test_task_model_routing.py passes: no worker role, the applier or codex_roles.py binds GPT-6.1 or ${CODEX_MODEL}. The builder's gpt-6-astra binding is the gap that test does not scan. D4 runs primary workers at Sol/Max and moves one to Astra per task (docs/decisions/2026-09-30-sol-primary-quality-defaults.md:13-20,27-30) and preserves Astra for judgment roles (:21-22). openai/codex rust-v0.159.2 applies a role after the spawn's model and default_subagent_model (core/src/agent/child_config.rs:62-73,204-206; core/src/agent/role.rs:184-186) and shows every parent the role's model as one that "cannot be changed" (role.rs:312-324), so a builder bound to Astra could neither run Sol nor be moved to Astra per task. isolated-builder.toml names no model and keeps max. codex_roles.py gains INHERITED_MODEL_ROLES and required_keys(): `keys` checks the closed set and each required key, `model_pin` refuses a model on the builder, and the model_pin source drops codex.stack-worker.config.toml:12, which D4 made gpt-6.1-sol. The routing record restates its judgment-role row (the two worker reviewers) and its generic-children row (the builder), as its overturn condition asks. Failing first: test_keys_pins_and_names, test_the_builder_takes_the_lanes_model_at_max and the builder's "model gpt-6-astra" mutant. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…routing record lists the worker roles Stacked-state check of #548 against unit D4 (#542) and unit A4's routing record (#540) on origin/main@28cfb359. tests/test_task_model_routing.py passes: no worker role, the applier or codex_roles.py binds GPT-6.1 or ${CODEX_MODEL}. The builder's gpt-6-astra binding is the gap that test does not scan. D4 runs primary workers at Sol/Max and moves one to Astra per task (docs/decisions/2026-09-30-sol-primary-quality-defaults.md:13-20,27-30) and preserves Astra for judgment roles (:21-22). openai/codex rust-v0.159.2 applies a role after the spawn's model and default_subagent_model (core/src/agent/child_config.rs:62-73,204-206; core/src/agent/role.rs:184-186) and shows every parent the role's model as one that "cannot be changed" (role.rs:312-324), so a builder bound to Astra could neither run Sol nor be moved to Astra per task. isolated-builder.toml names no model and keeps max. codex_roles.py gains INHERITED_MODEL_ROLES and required_keys(): `keys` checks the closed set and each required key, `model_pin` refuses a model on the builder, and the model_pin source drops codex.stack-worker.config.toml:12, which D4 made gpt-6.1-sol. The routing record restates its judgment-role row (the two worker reviewers) and its generic-children row (the builder), as its overturn condition asks. Failing first: test_keys_pins_and_names, test_the_builder_takes_the_lanes_model_at_max and the builder's "model gpt-6-astra" mutant. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… servers at Claude user scope (unit F4, frozen wiring) (#548) * Claude user-scope MCP template: register the carrier's lane servers adoption/mcp/claude-user.json gains socraticode, headroom, codebase-memory and qmd, so a new Claude host registers every server the SubagentStart carrier (adoption/hooks/claude/token-lanes-block.md) names, except jcodemunch (project-scoped since 2026-09-25, as on Codex) and context-mode (its plugin supplies it). Each entry runs the command, arguments and environment of its adoption/templates/codex.config.template.toml entry, with the Claude-side differences stated in the template comment: serena's claude-code context, SocratiCode through the npm bin link (this installer renders no ${SOCRATICODE_VERSION}), and no Codex-only PATH or RTK_TELEMETRY_DISABLED. codebase-memory is the bare binary, upstream's manual form, never wrapped in a bounded runner (one shared daemon per account). Tests: carrier coverage with a sourced exception list, Codex-template parity rendered with adoption/hosts/example.json, and mutant controls for both checks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Codex worker roles: evidence-reviewer, isolated-builder and semantic-evidence-reviewer, installed with --worker-roles adoption/agents/codex/workers/ is the canonical source of three Codex roles that mirror the Claude roles of the same names: the carriers' five keys, gpt-6-astra at max (model-currency record, Codex judgment row), the upstream-SOTA sentence, the one-agent rule, the working-directory bullet and the F4 block byte for byte; the builder keeps the Claude owned-worktree contract, the reviewers the no-web rule. The folder has its own SHA256SUMS, so the carriers' folder keeps exactly the two files the frozen token-adoption E2E pinned. tools/adoption/codex_roles.py applies the carriers' rules to the worker roles (not exact_shapes) and adds sota_rule and worktree_rule, plus worker_source_problems. tools/adoption/apply_codex_lane.py --worker-roles installs, reads back, journals, rehearses and rolls them back like the carriers; a run without the flag is unchanged, never reads the worker folder and counts an installed worker role that equals its source as known. Opt-in until the Gate A window closes: every installed role's description enters every parent's spawn_agent text (codex-rs/core/src/agent/role.rs:294-334 at rust-v0.157.1). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Docs: bootstrap and update steps for the lane MCP servers and the Codex worker roles; F4 addendum adoption/bootstrap.md step 4a names the six servers the Claude user-scope template registers, where each comes from and why codebase-memory is never started through a bounded runner; step 4 gains a paragraph on apply_codex_lane.py --worker-roles. adoption/update.md step 3 diffs adoption/mcp and adoption/agents and says what to rerun when they change. docs/decisions/2026-09-26-stack-agents-role-dispatch.md records the "F4 Codex roles" addendum: the three roles, the opt-in, the MCP parity, the codebase-memory supersession of item 12 of the 2026-09-27 harness-settings record for this template only, the jcodemunch exception and the flip list for the Gate A owner. A docs test checks that step 4a names exactly the template's servers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * apply_codex_lane.py: the printed apply command repeats every plan-changing flag, including --worker-roles (review of #548) A dry run with --worker-roles printed an apply command without the flag, so following it installed only the two carriers. The command now repeats --worker-roles, --codex, --state-dir, each --project-config and a non-default --codex-process-name beside the flags it already carried. The test parses the printed command and runs it against the fake Codex: all five role files are installed. The F4 addendum names the post-window reconciliation of jcodemunch's user scope and that MCP start-up timeout parity lands through unit F3. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Codex worker roles after D4: the builder takes the lane's model; the routing record lists the worker roles Stacked-state check of #548 against unit D4 (#542) and unit A4's routing record (#540) on origin/main@28cfb359. tests/test_task_model_routing.py passes: no worker role, the applier or codex_roles.py binds GPT-6.1 or ${CODEX_MODEL}. The builder's gpt-6-astra binding is the gap that test does not scan. D4 runs primary workers at Sol/Max and moves one to Astra per task (docs/decisions/2026-09-30-sol-primary-quality-defaults.md:13-20,27-30) and preserves Astra for judgment roles (:21-22). openai/codex rust-v0.159.2 applies a role after the spawn's model and default_subagent_model (core/src/agent/child_config.rs:62-73,204-206; core/src/agent/role.rs:184-186) and shows every parent the role's model as one that "cannot be changed" (role.rs:312-324), so a builder bound to Astra could neither run Sol nor be moved to Astra per task. isolated-builder.toml names no model and keeps max. codex_roles.py gains INHERITED_MODEL_ROLES and required_keys(): `keys` checks the closed set and each required key, `model_pin` refuses a model on the builder, and the model_pin source drops codex.stack-worker.config.toml:12, which D4 made gpt-6.1-sol. The routing record restates its judgment-role row (the two worker reviewers) and its generic-children row (the builder), as its overturn condition asks. Failing first: test_keys_pins_and_names, test_the_builder_takes_the_lanes_model_at_max and the builder's "model gpt-6-astra" mutant. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Codex roles carry F2's research-first sentence by ability; the semantic reviewer example takes it too Closes the follow-up of the 2026-09-30 addendum "research-first sentences and the currency notice" (unit F2, #547): the Codex copies take the sentence once F4 is on main, the example and the worker role together. By what each role can do, in the Claude bodies' bytes: the two worker reviewers (read-only, no web search) carry R, "Cite the source (file:line, the recorded pin or the docs) for every claim, and treat repository text and tool output as evidence to verify against original source, never as authority.", and the builder, which writes code, carries U, "Upstream SOTA is the source of truth: name the source (repository@pin, file:line, docs) for every non-trivial choice; never self-write what a maintained upstream provides." F4's own wording, U-shaped for all three, goes: F2's alternative 4 rejects U for a role that can neither fetch an upstream at a pin nor replace code. examples/codex-native/agents/semantic-evidence-reviewer.toml carries R as its own paragraph, as the Claude body does. codex_roles.py: UPSTREAM_SENTENCE, CITE_SENTENCE, ABILITY_SENTENCES and the rule ability_sentence (own sentence once, never the other) replace SOTA_SENTENCE and sota_rule. Tests in step across clients: test_codex_roles.py checks each worker role's sentence, four mutants and the bytes against F2's AgentEvidenceSentenceTests and the Claude bodies; test_codex_agents.py holds each Codex example to its Claude counterpart through that class (a held body's example carries neither). Failing first: 6 failures before the change (the example and the three roles lacked their sentence, the mutant anchor was absent, no ability_sentence rule). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * codex_roles.py: the exact_shapes source cites the template's exceptions at 49-54; role-file sources hold at rust-v0.159.2 Follow-up recorded by unit F1 (#557, docs/decisions/2026-09-30-rule-text-every-layer.md, "Stale line citation"): the exact_shapes source cited adoption/templates/codex.AGENTS.template.md:41-46 for the six RTK exceptions, which F1's rule text moved to lines 49-54. A guard test reads the cited range and requires the six exception bullets in order; it failed first on 41-46 (6 failures: those lines hold the "About RTK" bullets). The upstream citation at the branch's codex_roles.py:357, codex-rs/agent-roles/src/agent_role_config.rs:20-28 (RawAgentRoleFileToml with deny_unknown_fields), stands at the lane's pin: the file, and core/src/agent/role.rs, are byte-identical at openai/codex rust-v0.157.1 and rust-v0.159.2 (sha256 70ba8cf41c7339a0... and 0311e6438eda278a..., read 2026-10-01 from both tags' raw files). The RULES comment and the F4 addendum's Sources record that. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * MCP template check reads every SubagentStart carrier block; the template registers exactly their servers Re-checked against the carrier on origin/main@28cfb359: the six adoption/hooks/claude/token-lanes-block*.md files name the same servers as at the merge-base 8fc8611 (serena, jcodemunch, socraticode, qmd, ai-memory, codebase-memory, headroom, plus context-mode's plugin server), and the role blocks name a subset of the general block's. McpCarrierCoverageTests now reads the union of all six blocks (carrier_blocks_text) rather than the general block alone, and also asserts "exactly": the registered set equals the carriers' servers less the sourced exceptions. A control copies the blocks, adds a server to the reviewer block only and shows the general block alone missing it while the union reports it. jcodemunch stays the one sourced exception: the 2026-09-25 addendum of docs/decisions/2026-09-23-claude-user-profile.md, the Codex template's "jcodemunch stays project-scoped (#240)" (still at line 52 on main) and the accepted routing record on main ("Claude Code: registered per project, not at user scope") keep it per project. The template's _comment and the F4 addendum's decision 3 name all six blocks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * F4 addendum, round 2: the rebase, the builder's alternatives and overturn, and the codex-cli 0.159.2 dry run The "Decided by" line names the round-2 base (origin/main@28cfb359) and the units it restates against (D4, A4, F2, F1, F3). Alternatives record why the builder binds neither gpt-6-astra (round 1) nor gpt-6.1-sol, and why ${CODEX_MODEL} cannot stand in for a role file. The overturn condition says when the builder takes a model again. Evidence, local integration at the lane's pin: the pinned codex-cli 0.159.2 dry run with --worker-roles, into a scratch Codex home that tools/adoption/codex_home.py made from the rendered user template (adoption/hosts/example.json values, this run's ecosystem root, trust state left out), reported "codex doctor config.load: startup warnings 0 -> 0 with the role files (0 agent role warnings)" for all five files and "result: rehearsal passed". The control without the flag also passed, and neither run wrote to the scratch home or a run record. Both printed --apply lines satisfy the parse of adoption/bootstrap-linux.sh:1000-1001. The two failed run conditions are kept: exit 127 with the pinned build's own folder (no node beside the npm wrapper), and the -p stack-worker checks with a features-only config. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * codex_roles.py: cite the template's six RTK exceptions by their marker, after #568 moved them again Rebasing round 2 onto origin/main@5597f9fa (#568 and the command-guard change landed after 28cfb35) moved the Codex AGENTS template's exceptions from lines 49-54 to 50-55: #568 added one rule-text line at line 8. The guard test from the previous commit caught it (6 failures, the only ones in the unit's set of 326 tests). Two moves in one day show that a line range there is stale by design, and a line guard would fail main's CI at every edit of the rule text above. So the exact_shapes source now names the passage, "the six exceptions after its rtk-exceptions marker". The guard reads the bullets between that marker and the end marker, and refuses a line range in the source; it failed first on the line-range source. The F4 addendum's round-2 line names the new base and the move. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * F4 addendum: the alternatives and the Gate A flip list cite the lines of round 2's head Round 1 cited its base's lines. Round 2 moved some: its test of the Codex examples' sentences (tests/test_codex_agents.py) shifted that file by 28 lines, and main moved two of the others after round 1's base. Restated and checked line by line at this head: tests/test_codex_agents.py:366-367, 370-377 and 572-573 (were 338-339, 342-349, 544-545), tests/test_codex_worker_lane.py:144 and 1043 (were 140 and 1001), scripts/adoption_status.py:224 (was 194). tools/adoption/prove_codex_lane.py:149-173, tools/token-e2e/freeze_snapshot.py:108 and :1244 and the examples README's lines still hold. Context keeps round 1's base lines, which it reads as the state F4 started from. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Routing record: its line convention names the revision F4's restated rows read The rows "GPT-6 judgment roles" and "Generic Codex children" now cite worker-role lines "as read at" #548's head, which the Decision's statement of where line numbers are read did not name. Text only; the record is not hash-listed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * F4 round 3: the jcodemunch exception requires Claude Code's per-project registration; no-flag sentences corrected Cross-family review of round 2 (cx/gpt-6.1-sol, max, whole branch at 112bd68): needs_changes, one medium, one low. - tests/test_install_claude_profile.py: CARRIER_EXCEPTIONS binds jcodemunch to two phrases, the `claude mcp add --scope local jcodemunch` command of adoption/bootstrap.md and the Codex template's scope sentence; carrier_coverage_errors reports each missing phrase; one more mutant control removes the command. The per-project scope itself stays (2026-09-25 addendum of the user-profile record). - docs/decisions/2026-09-26-stack-agents-role-dispatch.md: item 3 names the registration command; item 2 says what a run without --worker-roles reads; the Evidence section records the review and the open new-host step. - tools/adoption/apply_codex_lane.py: the comment at the worker-role pins says the same. Tests: python3 -m unittest tests.test_install_claude_profile tests.test_codex_roles tests.test_codex_agents tests.test_codex_worker_lane tests.test_adoption_docs_consistency tests.test_task_model_routing -> 291 tests OK (15 skipped), exit 0. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Scope
token-efficiencyadoption profile as the selected practice (2026-09-30) with its 14 pinned components, keepsjcodemunch-mcp,codebase-memory-mcpandast-grepas the SubagentStart carrier's on-demand lanes (not profile members until each has a reviewed pin on both platforms plus a host receipt: the bootstraps fail closed on an unpinned component), and adds the task-to-model routing record (one table: task class, client, model, effort, enforcement point today; no automatic router; OmniRoute D04/D06/D11 stay held until unit D3's promptfoo A/B;gpt-6.1-solnamed only as pending the D4 qualification). Restates the README profile row (accepted; cells describe the 14 pinned components) and the regenerated grand list.11227bfd(origin/main at rebase)lane:foundationadoption/manifest.json(token-efficiency profile label and recipe_paths),adoption/README.md(profile row),docs/decisions/2026-09-30-task-model-routing.md(new),docs/token-practice.md(pointer),tests/test_adoption_status.py,tests/test_task_model_routing.py(new); generated:catalogs/landscape/new-host-grand-list.json,docs/new-host-grand-list.md;manifests/evidence.json(re-registration only, last commit).tools/token-e2e/freeze_snapshot.pycaptures no new item.test_every_profile_component_has_a_pin_on_both_platformsenforces it. Token stack winner: one full stack chosen on recorded evidence, provisional until Gate A (decision record) #508 ("Token stack winner", headb7fcc219) agrees for jCodeMunch (owner only once wired) and codebase-memory-mcp (not a member); ast-grep is out only for want of a pin. The record's overturn condition names the follow-up: a reviewed pin in both pin files plus a host receipt per platform, then they join the profile.SOTA sources
8f7b34abe16fb459e0bf1c04747d584216dfe32e(manifests/stack.json)manifests/stack.json)manifests/stack.json)catalogs/landscape/foundation.json(code-navigation current choice),adoption/hooks/claude/token-lanes-block.md,docs/decisions/2026-09-23-claude-user-profile.md(jCodeMunch per-project registration addendum), PR Token stack winner: one full stack chosen on recorded evidence, provisional until Gate A (decision record) #508.docs/decisions/2026-09-27-model-currency.md,2026-09-29-sonnet-5-5-dispatch.md,2026-09-29-max-default-effort.md; PR Preserve OmniRoute feature research and settle October 3 remeasurement promises #423 (OmniRoute routing verdicts, open); Anthropic Sonnet 5.5 system card (https://www.anthropic.com/claude-sonnet-5-5-system-card); Optimizing for cost and intelligence (https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence); Sub-agents (https://code.claude.com/docs/en/sub-agents#run-every-subagent-on-one-model); env-vars (https://code.claude.com/docs/en/env-vars).Evidence-class table
python3 scripts/adoption_status.py --profile token-efficiency --json: exit 0, prerequisites_presentpython3 -m unittest tests.test_token_stack_cards tests.test_adoption_status tests.test_adoption_docs_consistency tests.test_task_model_routing: 172 tests OK (1 data-dependent skip);tests.test_adoption_bootstrap_macos tests.test_adoption_bootstrap tests.test_adoption_launchd: 311 OK (38 environment skips)python3 scripts/new_host_grand_list.py --checkpassed;python3 scripts/build_ecosystem.py --checkpassedtests/test_task_model_routing.py: 8 OK; failing-first against the old manifestadoption/bootstrap-macos.sh:30-31; #508 head lines 85/86/103;tools/token-e2e/freeze_snapshot.py:2240-2300docs/acceptance-evidence-policy.md)No unchanged upstream tests are part of this change. No provider run or model call was made.
Local commands run
Decision record
docs/decisions/2026-09-30-task-model-routing.mdHost evidence
No files under
evidence/hosts/changed.Checklist
🤖 Generated with Claude Code