Repository navigation
Token layer gaps: routing rules, recipe limits and E1/E2 receipt errata from the 2026-09-27 verdict wave - #401
Merged
Conversation
…answer oracles Applies the 2026-09-27 cross-family GPT-6 review of the E1 (#296) and E2 (#316) receipts, append-only: every recorded field keeps its value, and each receipt keeps its own serialization. - Every tools[] row gains adjudication {as_of, task_acceptance, bad_input_control_recorded, basis, evidence}, using the review's status table and vocabulary (partial, retracted_later, untested, unsupported). - E1: QMD retracted_later (wrong_document, corrections_to_296); Repomix retracted_later (count=47 PASS on a 48-function file; repomix v1.18.1 PythonParseStrategy.ts L81-86 tests only a def's first line); ai-memory untested (base_empty=True is a vacuous pass); the other thirteen partial with no recorded failing control. A dated errata block names the pre-errata sha256 the preregistration records as a historical identity. - E2: context-mode, jcodemunch, qmd and ast-grep stay partial with no recorded failing control; SocratiCode's comparison counts an abridged transcription; RTK retracted by its verifier; Repomix's FAIL accurate. - READMEs carry dated corrections; the review is retained, sanitized, beside E1. - tests/test_token_e2e_receipt_checks.py: receipt contracts (red before the errata) and answer oracles rebuilt from pinned Git blobs, hash-matched to the recorded baselines, each with failing controls (red under a weaker historical-style check). Sources: docs/acceptance-evidence-policy.md (Discriminating controls); https://github.com/yamadashy/repomix/blob/v1.18.1/src/core/treeSitter/parseStrategies/PythonParseStrategy.ts#L81 evidence/artifacts/token-e2e-codex-20260926/receipt.json#/corrections_to_296. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Closes the documentation, recipe and routing gaps from the 2026-09-27 per-tool verdict wave for the token-efficiency layer. Every rule cites the pinned upstream source or a committed receipt; no pin changes. recipes/README.md (token-tool rows and sections): - agentsview: v0.44.0 unqualified; worker populations need --include-children, --include-automated and --include-one-shot, and --fts reads message bodies only (v0.43.0 docs/session-api.md, docs/commands.md); archive answers are observation. - ast-grep: every 0.45.3 command registers customLanguages from any sgconfig.yml in the working directory or a parent (lib.rs L107-139, config.rs L98-131/L280-299); PR #2960's opt-in is unreleased; outline (#2957) and YAML rule (#2963) limits. - ccusage: token-only while any model is unpriced (v20.0.24 json-output.md Unpriced Models; config-files.md pricingOverrides). - codebase-memory: trace_path excludes tests and evidence by default, search_code hits can be mentions (v0.11.0 src/mcp/mcp.c schemas); caller-set procedure. - context-hub: check metadata.versions (FastAPI guide 0.136.3 vs 0.141.1). - context-mode: npm tarball digest and plugin commits are separate identities. - headroom: v0.39.1 changes only the proxy limiter, so the omission-count defect stands. - markitdown: HTML as .txt passes through; -x html / -m text/html; selective extras. - mcporter: --name does not select the server; a positional token becomes the selector (v0.14.1 call-arguments.ts L164-192, ephemeral-flags.ts L97-105); the jCodeMunch example gains --server. - qmd: MCP query expands and reranks by default (v2.8.3 README, server.ts); typed lex + rerank:false + bounded get; catalog scope; refresh guard for #989/#991 (src/store.ts reindexCollection). - repomix: multi-line Python signatures are dropped (v1.18.1 PythonParseStrategy.ts L81-86); exact-definition tasks go to an uncompressed pack or the original. - toon: BOM root-string round trip (#339); no record-count threshold upstream; nested-uniform columns are tabular (tabular.ts L61-78). docs/token-session-handbook.md: a "Known upstream limits behind the lanes" section outside the carrier text (every carrier line unchanged), and row pointers. docs/token-practice.md: counts, comparisons and acceptance rules (invocation counts are not success rates; comparisons need task acceptance; count recovery reads; gpt-tokenizer's special-token contract; receipts carry adjudications). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ligibility wording
One cross-family review round (GPT-6 through codex-omniroute, read-only) returned
DEFECTS FOUND with two findings; both are fixed:
- P2 tests/test_token_e2e_receipt_checks.py: inventory_oracle checked only the claimed
count plus subset membership, so a correct count with missing, duplicated or no names
passed. It now requires the exact name set with one entry per function. New controls keep
the correct count while dropping, duplicating or inventing a name, or listing none; they
failed against the old oracle first (red) and pass now.
- P3 docs/token-session-handbook.md (and the matching recipes/README.md toon row): "accepts
any non-empty uniform array" overstated TOON 4.1.1's tabular eligibility. The encoder
needs non-empty objects sharing one key set, with columns of primitives or, recursively,
of non-empty objects sharing one key set; an empty object or an array-valued column falls
back to list form (packages/toon/src/encode/tabular.ts L6-78 at v4.1.1; confirmed with the
installed 4.1.1 CLI on [{}], [{"a":[]}], [{"a":{}}] and [{"a":{"b":1}}]).
Also: the ast-grep oracle keys counts by path under scripts/ rather than file stem, and E1's
Context Mode adjudication states that attempt 1's marker-grep FAIL on a refused call is not
shown to be a run of attempt 2's answer check (the receipt does not record that command).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… anchors, E2 Serena erratum The independent verifier of this branch reported five defects; all are fixed: - docs/token-session-handbook.md, the TOON Sources line under the carrier (on main before this branch) still said "The encoder accepts a non-empty uniform array", the wording GPT-6 finding P3 refuted. It now states the encoder's rule from packages/toon/src/encode/tabular.ts L6-78 at v4.1.1 and links the handbook's "Known upstream limits behind the lanes" section. It is not a "- " carrier line, so tests/test_token_lanes_subagent_start.py's verbatim-line check is unaffected. - Citation anchors: SocratiCode's INCLUDE_DOT_FILES default is documented under README.md "Indexing Behaviour" (v1.14.0 L1574-1579), not "Ignore Rules", so the handbook links #indexing-behaviour. codebase-memory-mcp v0.11.0 search_code gets its own link to src/mcp/mcp.c#L632-L659 in the handbook, and the recipe gives index_repository (L466-481), trace_path (L532-560) and search_code (L632-659) one link each. Both upstream files matched the retained copies by sha256 on 2026-09-27. - E2 receipt errata.items[4] and the E2 README erratum now name Serena's accurate reference FAIL (find_referencing_symbols missed 6 of 8 call sites; the definition check passed) and give each group of partial rows its own reason. Only that one finding string changed in receipt.json (same serialization, 1 line); written_at is unchanged because this refines the unmerged 2026-09-27 erratum on the same day. - tests/test_token_e2e_receipt_checks.py: the require_commit docstring now says the guard is modelled on RetainedEvidenceTests' Git check and adds the pinned-commit probe (tests/test_adoption_status.py L1721-1724 checks only git and the checkout). Sources: toon-format/toon v4.1.1 packages/toon/src/encode/tabular.ts; giancarloerra/SocratiCode v1.14.0 README.md; DeusData/codebase-memory-mcp v0.11.0 src/mcp/mcp.c; the E2 receipt's own serena row (tools[5]) and README. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
8 tasks done
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 27, 2026
manifests/evidence.json follows the hot-file protocol: main's copy, with this branch's files re-registered. #402 made .claude/agents/ a byte-for-byte project-scope copy of adoption/agents/claude/ (test_project_scope_agents_are_ the_installed_definitions), so landscape-sweep-worker.md gets its project copy. It is force-added and not hash-listed, like the other ten. Tests: test_install_claude_profile, test_token_lanes_subagent_start, test_adoption_docs_consistency, test_verdict_lane_vendoring, test_saturation_ledger, test_landscape_sweep_harness: 287 run, OK (7 skipped). validate.py passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes the token-efficiency layer's documentation, recipe, routing and receipt gaps from the 2026-09-27 per-tool verdict wave and from the GPT-6 review of the E1 (#296) and E2 (#316) subagent receipts. No pin, host configuration, Codex or Claude template, or OmniRoute record changes.
Recommended label:
lane:foundation.tools[]row of both receipts gains a datedadjudication(task_acceptance,bad_input_control_recorded,basis,evidence), and each receipt gains a top-levelerratablock that names its pre-errata sha256 (E1's is the identitypreregistration.jsonrecords inreceipt_sources). Recorded values are unchanged and each file keeps its own serialization (checked key by key).retracted_later(wrong document,corrections_to_296); Repomixretracted_later(count=47 PASSon a 48-function file); ai-memoryuntested(base_empty=True, a vacuous pass); the other thirteenpartial, since no E1 check recorded a failing run.partialwithbad_input_control_recorded: false; Serenapartial(its definition check passed and its reference check's FAIL is accurate); RTKretracted_later(verifier refutation); Repomixunsupported(accurate FAIL); SocratiCode's comparison is labelled an abridged transcription; Headroom's zero ledger is its own scope, not a measurement.evidence/artifacts/token-e2e-ultracode-20260925/review-e1e2-20260927.md.tests/test_token_e2e_receipt_checks.py:status_bodyandinspect), jCodeMunch location (E1 and E2), ast-grep call counts (E1 and E2, byte counts of the baselines), E1 Context Mode headings and E1 QMD answered document. Each rejects wrong answers. Swapping each oracle for the weaker historical-style check turns its control red (retained mutation run). These are 2026-09-27 local checks and upgrade no historical row.recipes/README.mdtoken-tool rows and sections, a new handbook section "Known upstream limits behind the lanes" (outside the carrier; every carrier line unchanged), and a token-practice section "Counts, comparisons and acceptance (2026-09-27)".SOTA sources
Every rule names the pinned upstream file or a committed receipt. Behaviour claims were checked against these sources on 2026-09-27, and all 37 newly cited GitHub URLs returned HTTP 200.
README.md#mcp-tool-parameters,src/mcp/server.tsL315-339 (thequerytool expands and reranks by default),src/store.tsL1605-1720 (reindexCollection), open issues #989 and #991.PythonParseStrategy.tsL81-86 (first-line signature regex).src/cli/call-arguments.tsL164-192 andsrc/cli/ephemeral-flags.tsL97-105. The installed 0.14.1 parser was run offline on the recipe's argv shapes: with--nameand no--server, akey=@filetoken becomes the selector and the arguments are empty;--serverfixes it.crates/cli/src/lib.rsL107-139 andconfig.rsL98-131 and L280-299 (config discovery and custom-language registration); PR #2960 (merged, unreleased); open #2957 and #2963.src/mcp/mcp.c:trace_pathL532-560 (include_testsandinclude_evidencedefaultfalse),search_codeL632-659 (a graph-ranked text search) and, in the recipe,index_repositoryL466-481.packages/toon/src/encode/tabular.tsL6-78 (no row minimum; a table needs non-empty objects sharing one key set, with primitive or, recursively, nested-uniform columns; an empty object or an array-valued column falls back to list form); open #339 (U+FEFF root string).docs/guide/json-output.md#unpriced-modelsanddocs/guide/config-files.md#pricing-overrides.docs/session-api.mdanddocs/commands.md(one-shot, automated and subagent sessions excluded by default;--ftssearches message bodies only); v0.44.0 release.__main__.pyL65-75 (-x/-mhints) and README optional dependencies. On a one-line, 102-byte synthetic HTML page saved as.txt(mdcheck/page.txt, the same bytes aspage.html), installed 0.1.8 passed the HTML through unchanged and converted it with-x htmlor-m text/html.encodethrows).6467ea36(versions: 0.136.3) against FastAPI 0.141.1.INCLUDE_DOT_FILESdefaultfalse, table row L1579).docs/acceptance-evidence-policy.md(Discriminating controls),docs/lanes.md(hot-file protocol), andevidence/artifacts/token-e2e-codex-20260926/receipt.json#/corrections_to_296.The retained copies of the TOON
tabular.ts, SocratiCodeREADME.mdand codebase-memorymcp.cinunits/tokgap/sources/matched the upstream raw files at the pinned tags by sha256 on 2026-09-27.Evidence classes
Acceptance
Run 2026-09-27T10:24:50Z at
fe60d3d6on base5f3a7c21(units/tokgap/acceptance.txt; full step logs inunits/tokgap/logs/):python3 scripts/validate.pypassed.bash rec/validate.sh python3printedFAILS=0(16 checks, includingevidence_manifest.py --check,component_matrix.py --checkandnew_host_grand_list.py --check).python3 scripts/build_ecosystem.py --checkpassed.python3 -m unittest tests.test_token_e2e_receipt_checks: 14 tests OK, none skipped.Pathjoin (units/tokgap/test-modules.txt): 2,026 tests OK (44 skipped).units/tokgap/apply_receipt_errata.py, applied to main's two receipts, regenerates both committed receipts byte for byte (E19332b722…, E28e15b963…).Cross-family review (GPT-6, one round)
codex-omniroute exec --skip-git-repo-check -s read-only(Codex CLI 0.157.1, modelcx/gpt-6-astrathrough the OmniRoute provider, reasoning effort max, approval never) overreview.diff, the whole patch before the review-fix commit (all ten files). 2026-09-27 08:28:37Z to 08:38:05Z, exit 0; Codex reported 110,101 tokens used. The prompt asked it to check every added claim against the cited pinned sources and the committed receipts, the append-only rule, the tests' soundness, the carrier lines and privacy.9bbb5807:tests/test_token_e2e_receipt_checks.py:123: the inventory oracle checked the claimed count plus subset membership, so a correct count with missing, duplicated or no names passed. It now requires the exact name set with one entry per function. New controls keep the correct count while dropping, duplicating or inventing a name, or listing none; three of them failed against the old oracle first (red-review-p2.txt), and the module's 14 tests pass after.docs/token-session-handbook.md:249: "accepts any non-empty uniform array" overstated TOON 4.1.1's tabular eligibility.9bbb5807put the encoder's rule fromtabular.tsL6-78 at v4.1.1 into the handbook's TOON limit and the recipes TOON row, confirmed with the installed 4.1.1 CLI on[{}],[{"a":[]}],[{"a":{}}]and[{"a":{"b":1}}](tooncheck/result.txt). The handbook's Sources line under the carrier (L228, on main before this branch) kept the refuted wording until the verifier round below.review.diff:9bbb5807also extended E1's Context Mode adjudication basis (attempt 1's marker-grep FAIL on a refused call is not shown to be a run of attempt 2's answer check) and re-keyed the ast-grep oracle by path underscripts/instead of file stem (subprocess_run_calls, nowtests/test_token_e2e_receipt_checks.pyL154-165). Its commit message names both; GPT-6 did not review them.Independent verifier (after the GPT-6 round)
Five findings, all fixed in
91812401. The manifest registration was then redone as the branch's only manifest commit (fe60d3d6, replacing4156aa0b). GPT-6 did not review these fixes.docs/token-session-handbook.md:228: the TOON Sources line still said "The encoder accepts a non-empty uniform array", which contradicted L249. It now states the encoder's rule, keeps "no five-record minimum" and links the "Known upstream limits" section. It is not a-carrier line, so the carrier's verbatim check is unaffected.INCLUDE_DOT_FILESdefault is documented under "Indexing Behaviour", not "Ignore Rules", and the anchor is corrected. codebase-memory'ssearch_codegets its own link (mcp.c#L632-L659, the whole schema entry) in the handbook and here, and the recipe givesindex_repository,trace_pathandsearch_codeone link each.errata.items[4]and the README erratum now name Serena's accurate reference FAIL (definition passed) and give each group of partial rows its own reason. Only that one finding string changed inreceipt.json. Itswritten_atstays08:09:36Z, because this refines the unmerged 2026-09-27 erratum on the same day, the same precedent9bbb5807set for E1's Context Mode basis.9bbb5807are named above, and mcporter#2 is marked closed in part, with its remaining copies listed.tests/test_token_e2e_receipt_checks.py:82: therequire_commitdocstring now says the guard is modelled on theRetainedEvidenceTestsGit check and adds the pinned-commit probe.Gap dispositions (44 assigned verdict gaps)
--serverwith--name), markitdown#1 (limitation stated), mcporter#1 (transport, not a saving), mcporter#2 (the recipe's jCodeMunch example gains--server, and the recipe asks copies to keep the complete form), qmd#1 (E1 adjudication, oracle and routing to a bounded read), qmd#3 (guard documented), repomix#2 (E1 adjudication and oracle), repomix#3 (routing, pack file list and exit status), serena#4 (language and path inclusion routing), toon#1 (upstream eligibility stated), toon#2 (keep the JSON unless the decode is value-equal); and the reading rule for invocation counts (token-practice) for ast-grep#3, ccusage#3, context-hub#3, context-mode#4, markitdown#3, rtk#4 and toon#3.50b9579f) puts rtk v0.50.0hooks/rtk-awareness-full.mdverbatim, plus this catalog's exceptions, inadoption/templates/codex.AGENTS.template.md, and the handbook's Codex paragraph records that Codex never expands@RTK.md.card-corrections.json): codebase-memory-mcp#3, jcodemunch-mcp#3, qmd#1, qmd#4, repomix#2, rtk#3, serena#4, socraticode#4, toon#1.Not done (with reasons)
install_pinfresh-host equivalence, each client's installed Context Mode plugin commit, and RTK's upstream Codex hook.evidence/artifacts/token-adoption-e2e-20260926, M1/M6c/M7/M8): success-confirmed, task-eligible use for ast-grep, ccusage, context-hub, context-mode (worker tool-access audit), headroom, markitdown, mcporter, repomix, rtk and toon.docs/token-practice.mdnow states how to read the existing invocation counts.units/tokgap/card-corrections.jsonlists each field, its current and corrected value and the retained source, for the assembler owner. No scratch-only figure is published in the docs.docs/token-efficiency-stack.md("All 16 tools worked") andobservability/grand-dashboard/state.json("16 of 16 tools"), the review's fifth finding. The corrected aggregate sentence is in the E1 errata for them to cite. No PR Token stack inside Ultracode subagents: 16/16 tools E2E with lifetime counters, exact comparisons and five new host receipts #296 addendum is posted from this unit.tools/token-report/README.mdL72-74, the argv it documents intools/token-report/token_manifest.pyL639-642, anddocs/token-efficiency-stack.jsonL722 and L1500 still pass--name jcodemunchwithout--server. They work, because their arguments go through--args: the offline 0.14.1 parser check parses that argv with no selector and the arguments set, and the retained 0.14.1 run of the same command (evidence/artifacts/sota-refresh-20260925/mcporter/results-x.json,S5-token-report-oneoff) exited 0 with the session stats. They do not use the complete--name/--server/--toolform the recipe now asks copies to keep; their owners should add--server jcodemunch.scripts/native_token_ci.py(the toon#2 fixture); a guarded refresh command inadoption/update.mdandcatalogs/us-equities/native-workflows.md(qmd#3; the guard is documented in the recipes and handbook); and the Codex and Claude template gaps.🤖 Generated with Claude Code