Skip to content

Token layer gaps: routing rules, recipe limits and E1/E2 receipt errata from the 2026-09-27 verdict wave - #401

Merged
seathatflowsinourveins merged 4 commits into
mainfrom
claude/token-layer-gaps-20260927
Sep 27, 2026
Merged

seathatflowsinourveins merged 4 commits into
mainfrom
claude/token-layer-gaps-20260927

Conversation

@seathatflowsinourveins

Copy link
Copy Markdown
Owner

Summary

Closes the token-efficiency layer's documentation, recipe, routing and receipt gaps from the 2026-09-27 per-tool verdict wave and from the GPT-6 review of the E1 (#296) and E2 (#316) subagent receipts. No pin, host configuration, Codex or Claude template, or OmniRoute record changes.

Recommended label: lane:foundation.

  1. Receipts, append-only. Every tools[] row of both receipts gains a dated adjudication (task_acceptance, bad_input_control_recorded, basis, evidence), and each receipt gains a top-level errata block that names its pre-errata sha256 (E1's is the identity preregistration.json records in receipt_sources). Recorded values are unchanged and each file keeps its own serialization (checked key by key).
    • E1: QMD retracted_later (wrong document, corrections_to_296); Repomix retracted_later (count=47 PASS on a 48-function file); ai-memory untested (base_empty=True, a vacuous pass); the other thirteen partial, since no E1 check recorded a failing run.
    • E2: Context Mode, jCodeMunch, QMD and ast-grep stay partial with bad_input_control_recorded: false; Serena partial (its definition check passed and its reference check's FAIL is accurate); RTK retracted_later (verifier refutation); Repomix unsupported (accurate FAIL); SocratiCode's comparison is labelled an abridged transcription; Headroom's zero ledger is its own scope, not a measurement.
    • Both READMEs carry dated corrections; the earlier review is retained, sanitized, as evidence/artifacts/token-e2e-ultracode-20260925/review-e1e2-20260927.md.
  2. Checker tests, red first. tests/test_token_e2e_receipt_checks.py:
    • Receipt contracts. They failed on the unmodified receipts (36 failures, 32 errors; retained red run) and pass after the errata.
    • Answer oracles rebuilt from pinned Git blobs, hash-matched to the recorded baselines: E1 Repomix inventory (48 defs; the first-line regex of repomix v1.18.1 predicts exactly status_body and inspect), jCodeMunch location (E1 and E2), ast-grep call counts (E1 and E2, byte counts of the baselines), E1 Context Mode headings and E1 QMD answered document. Each rejects wrong answers. Swapping each oracle for the weaker historical-style check turns its control red (retained mutation run). These are 2026-09-27 local checks and upgrade no historical row.
  3. Routing rules and recipe limits in recipes/README.md token-tool rows and sections, a new handbook section "Known upstream limits behind the lanes" (outside the carrier; every carrier line unchanged), and a token-practice section "Counts, comparisons and acceptance (2026-09-27)".

SOTA sources

Every rule names the pinned upstream file or a committed receipt. Behaviour claims were checked against these sources on 2026-09-27, and all 37 newly cited GitHub URLs returned HTTP 200.

The retained copies of the TOON tabular.ts, SocratiCode README.md and codebase-memory mcp.c in units/tokgap/sources/ matched the upstream raw files at the pinned tags by sha256 on 2026-09-27.

Evidence classes

  • Receipt adjudications: structural corrections of historical receipts; the recorded values are preserved.
  • Oracle tests: local integration checks against pinned Git content. They are not upstream tests, and not controls the historical runs recorded.
  • MCPorter parser and MarkItDown checks: local checks of the installed pinned tools on synthetic inputs, retained in the unit directory for this PR; nothing in the docs depends on them beyond the cited upstream source.
  • No provider run, model trial, host install or live GPU execution backs any claim here. The GPT-6 round and the independent verifier below reviewed the patch; they are not evidence for it.

Acceptance

Run 2026-09-27T10:24:50Z at fe60d3d6 on base 5f3a7c21 (units/tokgap/acceptance.txt; full step logs in units/tokgap/logs/):

  • python3 scripts/validate.py passed.
  • bash rec/validate.sh python3 printed FAILS=0 (16 checks, including evidence_manifest.py --check, component_matrix.py --check and new_host_grand_list.py --check).
  • python3 scripts/build_ecosystem.py --check passed.
  • python3 -m unittest tests.test_token_e2e_receipt_checks: 14 tests OK, none skipped.
  • The 40 unittest modules that reference a touched file, by literal path or a Path join (units/tokgap/test-modules.txt): 2,026 tests OK (44 skipped).
  • Red first: the receipt contracts failed on the unmodified receipts (36 failures, 32 errors); with each oracle swapped for its weaker historical-style check, every control goes red (inventory, symbol location, call counts, headings, release step) while the real oracles pass; and the review's new inventory controls failed against the old oracle (3 failures). All three runs are retained in the unit directory.
  • units/tokgap/apply_receipt_errata.py, applied to main's two receipts, regenerates both committed receipts byte for byte (E1 9332b722…, E2 8e15b963…).

Cross-family review (GPT-6, one round)

  • Run: codex-omniroute exec --skip-git-repo-check -s read-only (Codex CLI 0.157.1, model cx/gpt-6-astra through the OmniRoute provider, reasoning effort max, approval never) over review.diff, the whole patch before the review-fix commit (all ten files). 2026-09-27 08:28:37Z to 08:38:05Z, exit 0; Codex reported 110,101 tokens used. The prompt asked it to check every added claim against the cited pinned sources and the committed receipts, the append-only rule, the tests' soundness, the carrier lines and privacy.
  • Verdict: DEFECTS FOUND, two findings, both fixed in 9bbb5807:
    1. P2 tests/test_token_e2e_receipt_checks.py:123: the inventory oracle checked the claimed count plus subset membership, so a correct count with missing, duplicated or no names passed. It now requires the exact name set with one entry per function. New controls keep the correct count while dropping, duplicating or inventing a name, or listing none; three of them failed against the old oracle first (red-review-p2.txt), and the module's 14 tests pass after.
    2. P3 docs/token-session-handbook.md:249: "accepts any non-empty uniform array" overstated TOON 4.1.1's tabular eligibility. 9bbb5807 put the encoder's rule from tabular.ts L6-78 at v4.1.1 into the handbook's TOON limit and the recipes TOON row, confirmed with the installed 4.1.1 CLI on [{}], [{"a":[]}], [{"a":{}}] and [{"a":{"b":1}}] (tooncheck/result.txt). The handbook's Sources line under the carrier (L228, on main before this branch) kept the refuted wording until the verifier round below.
  • Not in review.diff: 9bbb5807 also extended E1's Context Mode adjudication basis (attempt 1's marker-grep FAIL on a refused call is not shown to be a run of attempt 2's answer check) and re-keyed the ast-grep oracle by path under scripts/ instead of file stem (subprocess_run_calls, now tests/test_token_e2e_receipt_checks.py L154-165). Its commit message names both; GPT-6 did not review them.
  • Residuals: none. One round only, as the unit brief requires; no second GPT-6 run.

Independent verifier (after the GPT-6 round)

Five findings, all fixed in 91812401. The manifest registration was then redone as the branch's only manifest commit (fe60d3d6, replacing 4156aa0b). GPT-6 did not review these fixes.

  1. Medium, docs/token-session-handbook.md:228: the TOON Sources line still said "The encoder accepts a non-empty uniform array", which contradicted L249. It now states the encoder's rule, keeps "no five-record minimum" and links the "Known upstream limits" section. It is not a - carrier line, so the carrier's verbatim check is unaffected.
  2. Low, citation anchors: SocratiCode's INCLUDE_DOT_FILES default is documented under "Indexing Behaviour", not "Ignore Rules", and the anchor is corrected. codebase-memory's search_code gets its own link (mcp.c#L632-L659, the whole schema entry) in the handbook and here, and the recipe gives index_repository, trace_path and search_code one link each.
  3. Low, E2 aggregate erratum: errata.items[4] and the README erratum now name Serena's accurate reference FAIL (definition passed) and give each group of partial rows its own reason. Only that one finding string changed in receipt.json. Its written_at stays 08:09:36Z, because this refines the unmerged 2026-09-27 erratum on the same day, the same precedent 9bbb5807 set for E1's Context Mode basis.
  4. Low, this PR body: the measurement list below now matches the dispositions, the MarkItDown input is described accurately, the unreviewed extras in 9bbb5807 are named above, and mcporter#2 is marked closed in part, with its remaining copies listed.
  5. Low, tests/test_token_e2e_receipt_checks.py:82: the require_commit docstring now says the guard is modelled on the RetainedEvidenceTests Git check and adds the pinned-commit probe.

Gap dispositions (44 assigned verdict gaps)

  • Closed here: agentsview#2 and Add searchable ecosystem manifest and token-efficiency guide #3, ast-grep#2, ccusage#2, codebase-memory-mcp#1, context-hub#1 and Publish evidence-led convergence practice and retrieval comparisons #2, gpt-tokenizer#3, headroom#3, markitdown#2, qmd#2, repomix#1, rtk#2.
  • Closed here in part, the rest listed below: ast-grep#1 (config-scope routing), ccusage#1 (token-only rule), context-mode#3 (npm digest and plugin commit recorded as separate identities), gpt-tokenizer#2 (counting-contract boundary), headroom#1 (0.39.1 does not touch the defect), headroom#4 (--server with --name), markitdown#1 (limitation stated), mcporter#1 (transport, not a saving), mcporter#2 (the recipe's jCodeMunch example gains --server, and the recipe asks copies to keep the complete form), qmd#1 (E1 adjudication, oracle and routing to a bounded read), qmd#3 (guard documented), repomix#2 (E1 adjudication and oracle), repomix#3 (routing, pack file list and exit status), serena#4 (language and path inclusion routing), toon#1 (upstream eligibility stated), toon#2 (keep the JSON unless the decode is value-equal); and the reading rule for invocation counts (token-practice) for ast-grep#3, ccusage#3, context-hub#3, context-mode#4, markitdown#3, rtk#4 and toon#3.
  • Already on main: rtk#1's instruction half. PR-D: Codex worker lane: Codex's own config writer, RTK text inline with exceptions, max-effort stack-worker profile #389 (50b9579f) puts rtk v0.50.0 hooks/rtk-awareness-full.md verbatim, plus this catalog's exceptions, in adoption/templates/codex.AGENTS.template.md, and the handbook's Codex paragraph records that Codex never expands @RTK.md.
  • Not done, pin or host qualification: agentsview#1, ast-grep#1, ccusage#1 (qualified prices), context-mode#3 (reading each client's installed plugin commit on a host), gpt-tokenizer#1, headroom#1, markitdown#1, rtk#1 (upstream Codex hook).
  • Not done, measurement: ast-grep#3, ccusage#3, context-hub#3, context-mode#4, headroom#4, markitdown#3, mcporter#1, repomix#3, rtk#3, rtk#4, toon#3.
  • Not done, card data (handed off in card-corrections.json): codebase-memory-mcp#3, jcodemunch-mcp#3, qmd#1, qmd#4, repomix#2, rtk#3, serena#4, socraticode#4, toon#1.
  • Not done, outside this unit's paths: qmd#3, toon#2, gpt-tokenizer#2, mcporter#2 (the remaining command copies; see below).

Not done (with reasons)

  • Pin upgrades needing host installs and qualification: agentsview 0.44.0, gpt-tokenizer 4.0.0, a headroom release past the 0.37.0 pin (the omission-count defect stands at 0.39.1), a released ast-grep with PR #2960, qualified ccusage prices for the unpriced models, the MarkItDown bootstrap install_pin fresh-host equivalence, each client's installed Context Mode plugin commit, and RTK's upstream Codex hook.
  • Measurement gaps that need the preregistered E2E (evidence/artifacts/token-adoption-e2e-20260926, M1/M6c/M7/M8): success-confirmed, task-eligible use for ast-grep, ccusage, context-hub, context-mode (worker tool-access audit), headroom, markitdown, mcporter, repomix, rtk and toon. docs/token-practice.md now states how to read the existing invocation counts.
  • Card-data corrections (codebase-memory-mcp#3, jcodemunch-mcp#3, serena#4 and socraticode#4 comparison lists; qmd#4 count and comparisons; the card halves of qmd#1 and repomix#2; rtk#3's PR-D label; toon#1's shortfall text): the cards are scratch assembler output outside this unit's paths. units/tokgap/card-corrections.json lists each field, its current and corrected value and the retained source, for the assembler owner. No scratch-only figure is published in the docs.
  • Explorer-owned aggregate text: docs/token-efficiency-stack.md ("All 16 tools worked") and observability/grand-dashboard/state.json ("16 of 16 tools"), the review's fifth finding. The corrected aggregate sentence is in the E1 errata for them to cite. No PR Token stack inside Ultracode subagents: 16/16 tools E2E with lifetime counters, exact comparisons and five new host receipts #296 addendum is posted from this unit.
  • MCPorter command copies outside this unit's paths (mcporter#2): tools/token-report/README.md L72-74, the argv it documents in tools/token-report/token_manifest.py L639-642, and docs/token-efficiency-stack.json L722 and L1500 still pass --name jcodemunch without --server. They work, because their arguments go through --args: the offline 0.14.1 parser check parses that argv with no selector and the arguments set, and the retained 0.14.1 run of the same command (evidence/artifacts/sota-refresh-20260925/mcporter/results-x.json, S5-token-report-oneoff) exited 0 with the session stats. They do not use the complete --name/--server/--tool form the recipe now asks copies to keep; their owners should add --server jcodemunch.
  • Other owned-elsewhere changes: adding TOON's BOM case to scripts/native_token_ci.py (the toon#2 fixture); a guarded refresh command in adoption/update.md and catalogs/us-equities/native-workflows.md (qmd#3; the guard is documented in the recipes and handbook); and the Codex and Claude template gaps.
  • Retaining gpt-tokenizer's special-token failure cases as a committed receipt: the cases are in scratch only, so the docs state the contract boundary from the upstream README instead.
  • Complete native outputs for E1 and E2 (the earlier review's proposal 7): both runs kept them only by sha256 in the host's private ledger. The errata say so, and the oracles rebuild what pinned Git content can reproduce.

🤖 Generated with Claude Code

Scout and others added 4 commits September 27, 2026 09:22
…answer oracles

Applies the 2026-09-27 cross-family GPT-6 review of the E1 (#296) and E2 (#316)
receipts, append-only: every recorded field keeps its value, and each receipt keeps
its own serialization.

- Every tools[] row gains adjudication {as_of, task_acceptance, bad_input_control_recorded,
  basis, evidence}, using the review's status table and vocabulary (partial,
  retracted_later, untested, unsupported).
- E1: QMD retracted_later (wrong_document, corrections_to_296); Repomix retracted_later
  (count=47 PASS on a 48-function file; repomix v1.18.1 PythonParseStrategy.ts L81-86
  tests only a def's first line); ai-memory untested (base_empty=True is a vacuous pass);
  the other thirteen partial with no recorded failing control. A dated errata block
  names the pre-errata sha256 the preregistration records as a historical identity.
- E2: context-mode, jcodemunch, qmd and ast-grep stay partial with no recorded failing
  control; SocratiCode's comparison counts an abridged transcription; RTK retracted by
  its verifier; Repomix's FAIL accurate.
- READMEs carry dated corrections; the review is retained, sanitized, beside E1.
- tests/test_token_e2e_receipt_checks.py: receipt contracts (red before the errata) and
  answer oracles rebuilt from pinned Git blobs, hash-matched to the recorded baselines,
  each with failing controls (red under a weaker historical-style check).

Sources: docs/acceptance-evidence-policy.md (Discriminating controls);
https://github.com/yamadashy/repomix/blob/v1.18.1/src/core/treeSitter/parseStrategies/PythonParseStrategy.ts#L81
evidence/artifacts/token-e2e-codex-20260926/receipt.json#/corrections_to_296.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Closes the documentation, recipe and routing gaps from the 2026-09-27 per-tool
verdict wave for the token-efficiency layer. Every rule cites the pinned upstream
source or a committed receipt; no pin changes.

recipes/README.md (token-tool rows and sections):
- agentsview: v0.44.0 unqualified; worker populations need --include-children,
  --include-automated and --include-one-shot, and --fts reads message bodies only
  (v0.43.0 docs/session-api.md, docs/commands.md); archive answers are observation.
- ast-grep: every 0.45.3 command registers customLanguages from any sgconfig.yml in
  the working directory or a parent (lib.rs L107-139, config.rs L98-131/L280-299);
  PR #2960's opt-in is unreleased; outline (#2957) and YAML rule (#2963) limits.
- ccusage: token-only while any model is unpriced (v20.0.24 json-output.md
  Unpriced Models; config-files.md pricingOverrides).
- codebase-memory: trace_path excludes tests and evidence by default, search_code
  hits can be mentions (v0.11.0 src/mcp/mcp.c schemas); caller-set procedure.
- context-hub: check metadata.versions (FastAPI guide 0.136.3 vs 0.141.1).
- context-mode: npm tarball digest and plugin commits are separate identities.
- headroom: v0.39.1 changes only the proxy limiter, so the omission-count defect stands.
- markitdown: HTML as .txt passes through; -x html / -m text/html; selective extras.
- mcporter: --name does not select the server; a positional token becomes the
  selector (v0.14.1 call-arguments.ts L164-192, ephemeral-flags.ts L97-105); the
  jCodeMunch example gains --server.
- qmd: MCP query expands and reranks by default (v2.8.3 README, server.ts); typed
  lex + rerank:false + bounded get; catalog scope; refresh guard for #989/#991
  (src/store.ts reindexCollection).
- repomix: multi-line Python signatures are dropped (v1.18.1 PythonParseStrategy.ts
  L81-86); exact-definition tasks go to an uncompressed pack or the original.
- toon: BOM root-string round trip (#339); no record-count threshold upstream;
  nested-uniform columns are tabular (tabular.ts L61-78).

docs/token-session-handbook.md: a "Known upstream limits behind the lanes" section
outside the carrier text (every carrier line unchanged), and row pointers.
docs/token-practice.md: counts, comparisons and acceptance rules (invocation counts
are not success rates; comparisons need task acceptance; count recovery reads;
gpt-tokenizer's special-token contract; receipts carry adjudications).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ligibility wording

One cross-family review round (GPT-6 through codex-omniroute, read-only) returned
DEFECTS FOUND with two findings; both are fixed:

- P2 tests/test_token_e2e_receipt_checks.py: inventory_oracle checked only the claimed
  count plus subset membership, so a correct count with missing, duplicated or no names
  passed. It now requires the exact name set with one entry per function. New controls keep
  the correct count while dropping, duplicating or inventing a name, or listing none; they
  failed against the old oracle first (red) and pass now.
- P3 docs/token-session-handbook.md (and the matching recipes/README.md toon row): "accepts
  any non-empty uniform array" overstated TOON 4.1.1's tabular eligibility. The encoder
  needs non-empty objects sharing one key set, with columns of primitives or, recursively,
  of non-empty objects sharing one key set; an empty object or an array-valued column falls
  back to list form (packages/toon/src/encode/tabular.ts L6-78 at v4.1.1; confirmed with the
  installed 4.1.1 CLI on [{}], [{"a":[]}], [{"a":{}}] and [{"a":{"b":1}}]).

Also: the ast-grep oracle keys counts by path under scripts/ rather than file stem, and E1's
Context Mode adjudication states that attempt 1's marker-grep FAIL on a refused call is not
shown to be a run of attempt 2's answer check (the receipt does not record that command).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… anchors, E2 Serena erratum

The independent verifier of this branch reported five defects; all are fixed:

- docs/token-session-handbook.md, the TOON Sources line under the carrier (on main
  before this branch) still said "The encoder accepts a non-empty uniform array",
  the wording GPT-6 finding P3 refuted. It now states the encoder's rule from
  packages/toon/src/encode/tabular.ts L6-78 at v4.1.1 and links the handbook's
  "Known upstream limits behind the lanes" section. It is not a "- " carrier line,
  so tests/test_token_lanes_subagent_start.py's verbatim-line check is unaffected.
- Citation anchors: SocratiCode's INCLUDE_DOT_FILES default is documented under
  README.md "Indexing Behaviour" (v1.14.0 L1574-1579), not "Ignore Rules", so the
  handbook links #indexing-behaviour. codebase-memory-mcp v0.11.0 search_code gets
  its own link to src/mcp/mcp.c#L632-L659 in the handbook, and the recipe gives
  index_repository (L466-481), trace_path (L532-560) and search_code (L632-659)
  one link each. Both upstream files matched the retained copies by sha256 on
  2026-09-27.
- E2 receipt errata.items[4] and the E2 README erratum now name Serena's accurate
  reference FAIL (find_referencing_symbols missed 6 of 8 call sites; the definition
  check passed) and give each group of partial rows its own reason. Only that one
  finding string changed in receipt.json (same serialization, 1 line); written_at
  is unchanged because this refines the unmerged 2026-09-27 erratum on the same day.
- tests/test_token_e2e_receipt_checks.py: the require_commit docstring now says the
  guard is modelled on RetainedEvidenceTests' Git check and adds the pinned-commit
  probe (tests/test_adoption_status.py L1721-1724 checks only git and the checkout).

Sources: toon-format/toon v4.1.1 packages/toon/src/encode/tabular.ts;
giancarloerra/SocratiCode v1.14.0 README.md; DeusData/codebase-memory-mcp v0.11.0
src/mcp/mcp.c; the E2 receipt's own serena row (tools[5]) and README.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Sep 27, 2026
@seathatflowsinourveins
seathatflowsinourveins merged commit 5b72d41 into main Sep 27, 2026
24 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the claude/token-layer-gaps-20260927 branch September 27, 2026 13:43
seathatflowsinourveins pushed a commit that referenced this pull request Sep 27, 2026
manifests/evidence.json follows the hot-file protocol: main's copy, with
this branch's files re-registered. #402 made .claude/agents/ a byte-for-byte
project-scope copy of adoption/agents/claude/ (test_project_scope_agents_are_
the_installed_definitions), so landscape-sweep-worker.md gets its project
copy. It is force-added and not hash-listed, like the other ten.
Tests: test_install_claude_profile, test_token_lanes_subagent_start,
test_adoption_docs_consistency, test_verdict_lane_vendoring,
test_saturation_ledger, test_landscape_sweep_harness: 287 run, OK
(7 skipped). validate.py passed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant