Skip to content

Token stack inside Ultracode subagents: 16/16 tools E2E with lifetime counters, exact comparisons and five new host receipts - #296

Merged
seathatflowsinourveins merged 3 commits into
mainfrom
claude/token-e2e-ultracode-20260925
Sep 25, 2026
Merged

seathatflowsinourveins merged 3 commits into
mainfrom
claude/token-e2e-ultracode-20260925

Conversation

@seathatflowsinourveins

Copy link
Copy Markdown
Owner

Scope

  • What this PR changes: shows that each selected token-efficiency tool runs inside Ultracode workflow subagents on this workstation, and that its savings show up in the tool's own lifetime counters and in exact tokenizer comparisons. It also records install and use receipts for the five tools installed today.
  • Base commit: c303d8e3 (current origin/main, after Codex MCP servers at user scope so every worktree inherits the token tools; --pinned-versions coverage check #289)
  • Lane: lane:foundation
  • Owned paths touched:
    • evidence/artifacts/token-e2e-ultracode-20260925/ (README and receipt.json)
    • 11 receipts under evidence/hosts/nativestack-5975wx-20260925/ (ast-grep, codebase-memory-mcp, context-hub, agentsview, otel-tui)
    • a new section in docs/token-efficiency-stack.md
    • regenerated matrix and grand list, and manifests/evidence.json registrations

Why. Installing and configuring a tool does not show a saving. Each claim here pairs a subagent actually using the tool, with its answer checked against the plain baseline, with that tool's own counter where it has one and an exact o200k_base comparison of the same information need.

Host changes made before the run

These are not part of this diff.

  • Installed at their manifests/stack.json pins, checked against upstream sha256 or npm integrity: ast-grep 0.45.3, codebase-memory-mcp 0.11.0, Context Hub 0.1.4, agentsview 0.43.0, otel-tui 0.7.5.
  • Registered as Claude user-scope MCP servers: headroom (headroom mcp serve --proxy-url http://127.0.0.1:1, so no model traffic is rerouted), codebase-memory (with its web UI turned off: ui_enabled=false) and qmd (qmd --index native-agent-stack-catalog mcp). Codex MCP servers at user scope so every worktree inherits the token tools; --pinned-versions coverage check #289 mirrors these in Codex.
  • Ledger: the private token-report ledger now reads Context Mode stats for both clients and uses gpt-tokenizer@3.4.0.

Evidence-class table

Claim Evidence class Where
16/16 tools ran inside a Sonnet 5 workflow subagent, and each subagent's answer passed a deterministic check against the baseline. Context Mode, Serena and Context Hub passed on a second attempt, after a real constraint. native_proven on this host receipt.json tools[]
Exact tokens kept out of subagent context, per task: 12.5% (codebase-memory) to 99.8% (Context Mode) local_integration (exact o200k_base counts of retained artifacts; not provider savings) tools[].exact_comparison
RTK counted +56,527 saved for this run's worktree alone (89 commands), isolated by working directory upstream estimate counters.rtk_gain_e2e_worktree_only
Headroom's lifetime counter moved +183,904; the exact comparison of the same compression measured 180,752 upstream estimate plus exact counters.headroom_savings_lifetime
Subagent cost: 19 children, 649,193 output tokens and 13.9M cache-read tokens provider-returned, per child workflow_consumption (child-usage.mjs)
A fresh claude -p session loads the three new servers and calls one tool on each native_proven fresh_native_session
Install and use receipts for the five new tools native_proven (self review only) evidence/hosts/…

Retained failures and gaps

The full list is in the README.

  • Codex workers were not run: codex exec returns the usage limit until 2026-09-30 23:50.
  • The jCodeMunch persistent ledger read +68, while the live session reported 23,643 not yet written.
  • The secret guard flagged a harmless heredoc set(...) as environment_dump (a false positive).
  • MarkItDown needs an .html extension; on .txt it passes the input through.
  • OmniRoute was not exercised, because it reroutes model traffic and needs a credential.
  • Four tools run newer versions than their pins after another session's upgrade: rtk 0.50.0, markitdown 0.1.8, ai-memory 2.4.0, mcporter 0.14.1.

Local commands run

$ uv run … scripts/validate.py; evidence_manifest.py --check; build_ecosystem.py --check   # passed (5625 hashed files)
$ uv run … scripts/component_matrix.py --check; new_host_grand_list.py --check; host_receipts.py validate   # checked / passed / 109 receipts passed
$ uv run … scripts/verdict_review_gate.py --base origin/main   # passed (builder run)
$ uv run … -m unittest <every tests/ module that reads token-efficiency-stack or evidence/artifacts>   # Ran 1391, OK (skipped=24)
$ gitleaks-guarded dir evidence/artifacts/token-e2e-ultracode-20260925   # no leaks found; the pre-commit hook passed

uv run … stands for uv run --no-project --python 3.13 --with-requirements .github/requirements-ci.txt python.

Decision record

None. This PR is evidence only and makes no pin or verdict change.

Host evidence

  • host_receipts.py validate passes.
  • Independent review of the 11 receipts is requested. A separate headless Claude session will record independent_session verdicts. Codex is unavailable until 2026-09-30.
  • No platform_status change is made by hand.

Checklist

  • No workflow changes. [x] No secrets. The receipt was scanned for home paths, scratch paths and session ids: 0 hits. [x] No paid surface. [x] Peer files preserved.

🤖 Generated with Claude Code

Scout and others added 2 commits September 25, 2026 17:54
…-5975wx-20260925

Install and use-stage receipts for ast-grep 0.45.3, codebase-memory-mcp 0.11.0,
context-hub 0.1.4, agentsview 0.43.0, and otel-tui 0.7.5, each pinned to their
manifests/stack.json versions and recorded with scripts/host_receipts.py per
docs/contributing-evidence.md. Each use-stage command asserts its own result
in-line (a positive control plus a cheap negative control) so the printed
output backs the claim without relying on a human to eyeball a clean run.
The codebase-memory-mcp use receipt was superseded once to keep its assertion
lines inside the 400-char excerpt window; the original stays byte-identical.
Refreshes the derived component-evidence-matrix and evidence manifest.

All receipts carry only the recorder's self review; independent review is
still required before any of these count toward accepted status.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ime counters

Sixteen Sonnet 5 workflow subagents each used one selected tool on real work
in this repository and checked the answer against the plain baseline (all 16
pass; Context Mode, Serena and Context Hub on a second attempt after a real
binding constraint). Upstream counters moved over the 28-minute window: RTK
+56,527 saved for this run's worktree alone and Headroom +183,904, which the
exact o200k comparison of the same compression matches at 180,752. Thirteen
exact per-task comparisons, per-child provider usage, a fresh-session MCP
check and the retained failures are in the receipt. Codex workers were not
run: the account is at its usage limit until 2026-09-30.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Sep 25, 2026
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ⚠️ Failed 2026-09-25T21:56:11.260189Z 0af4e7f PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

… (PR #296)

A separate headless Claude session (Sonnet) recorded agree on all 11
receipts (ast-grep, codebase-memory-mcp, context-hub, agentsview, otel-tui
install/use). It found one mismatch in the E2E README: the fresh-session
server list omitted the claude.ai Claude Docs connector, now listed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins merged commit d0d8c5b into main Sep 25, 2026
23 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the claude/token-e2e-ultracode-20260925 branch September 25, 2026 22:40
seathatflowsinourveins added a commit that referenced this pull request Sep 25, 2026
… request lane (#298)

Two foundation gates. token-stack-subagent-e2e points at #296's receipt: 16 of
16 tools used in Sonnet workflow subagents with passing checks, RTK +56527
for the run-only worktree, Headroom +183904 lifetime, Codex workers blocked
until 2026-09-30. host-request-lane points at the #266 recipe: the workstation
poll timer runs every 10 min, 0 workstation requests, 2 mac-coordinator
requests open (#274, #276). Values use the dashboard token alphabet.

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 25, 2026
Keeps the branch current with main after #296, #297, #291 and #298 so CI
judges it against the current base.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 25, 2026
… DO_NOT_TRACK=1) (#304)

headroom-ai 0.37.0 uploads an anonymous usage beacon by default
(telemetry/beacon.py BEACON_DEFAULT_ON = True; ccr/mcp_server.py calls it on
every MCP compression). offline.py's HEADROOM_OFFLINE is the master no-egress
switch (beacon, update check, usage reporter, HF downloads); DO_NOT_TRACK also
disables the beacon. The recipe and token-stack row now set both, with the
native Claude registration command; the #296 E2E README records that the
beacon was at its default during that run.

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Upstream check of two #296 registrations (from the laptop lane, wsl-authoring-20260923; upstream is the source of truth)

  1. Headroom MCP registered on its own.
    • Headroom's MCP docs (https://docs.headroomlabs.ai/docs/mcp, "MCP + Proxy (full setup)") say: "The proxy compresses all traffic at the HTTP level (before the LLM sees content). MCP tools operate after the LLM receives content."
    • Their troubleshooting section recommends: "Prefer the proxy path for automatic compression of normal Claude Code traffic."
    • Token stack inside Ultracode subagents: 16/16 tools E2E with lifetime counters, exact comparisons and five new host receipts #296's 237,691 → 56,939 came from piping git log output through MCPorter in Bash, so the content never entered the model's context. A native headroom_compress(content) call needs the model to re-emit that content first.
    • So the MCP registration by itself is not a token saver for normal traffic. The saver upstream means is the proxy (headroom wrap claude or headroom install apply --preset persistent-service).
    • That proxy also changes model behaviour: "Effort routing dials thinking effort down when a turn is only the model resuming after a tool result" (README v0.39.0). It conflicts with the max-effort worker rule, so it needs the measured comparison before adoption.
  2. codebase-memory-mcp registered at user scope.

The laptop does not register either. Please reconcile the workstation's user-scope set against these upstream statements and the catalog's disposition.

seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
…acts and receipt)

Qualified on nativestack-5975wx-20260925 against the pinned 0.37.0: the
verified wheel in a new prefix (RECORD 555/555), the unchanged upstream
#3736 test 6/6 (0.37.0: 4/6), and the stack's MCP path over stdio with
byte-exact retrieval and cross-version store reads; an independent
verifier agreed. The coordinator retains 0.37.0: at v0.39.0 the
'[N lines omitted: ...]' notice counts levels over all input lines
(log_compressor.py:461-487 and the Rust format_output; checked with gh),
so fixture A's '[593 lines omitted: 1 ERROR]' reports the kept CRITICAL
line as omitted; the Rust detector still lacks the #3736 timestamp
guard; #299 pins 0.37.0 in the bootstrap and #296's counters run at it.
No pin changes.

- evidence/artifacts/sota-refresh-20260925/headroom-039/: harness and
  verifier scripts, key results as files, all other outputs and the
  production snapshots bundled, a read-only read-back and results.json
  (decision reasons, overturn condition, beacon note, corrections).
- evidence/receipts/headroom-039-qualification-20260925.json
  (native_cli_e2e).

manifests/evidence.json registration follows in the branch's last commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
…station SOTA refresh

docs/decisions/2026-09-25-ai-memory-2-4-0-release-review.md gains a dated
correction section; the original text is kept. GHSA-9pj6-vhgr-3mwh is not
reachable on this host (stateless transport: GET /mcp 405 with allow:
POST, no --http-stateful), and 2.4.0 writes no automatic pre-migration
archive for a store created by 2.3.2 (snapshot_before_db_migration is
gated to the 1.x -> 2.0 upgrade). It also records that the upgrade was
performed on 2026-09-25 and cites the new receipt.

docs/decisions/2026-09-25-workstation-sota-refresh.md is the dated
research-convergence record for the whole workstation-lane refresh: per
unit, the decision, its primary sources (releases, compares, advisories
and the refuter's key finding), the alternatives and the overturn
condition. rtk, markitdown, mcporter and ai-memory were qualified and
switched; socraticode 1.15.0 is qualified with the cutover deferred to no
earlier than 2026-10-01; dagu 2.17.2 is held by the trading lane;
headroom stays 0.37.0 because 0.38.0 carries #3736 on the MCP compress
path; qmd, qdrant, ccusage, worktrunk and vllm are already current;
llama.cpp b11146 is recorded, not qualified; both models are retained;
the ai-memory Nemotron embedder is evaluation-only. The research and
qualification results are cited as private inputs only.

Review repairs (2026-09-25, independent review needs_changes), each
fact re-checked with gh api at about 22:35Z:
- headroom: v0.39.0 (2026-09-25T19:29:55Z) contains d971f7c3 (compare
  ahead 86, behind 0; the release notes list #3748), so the overturn
  condition is met; qualifying 0.39.0 is the queued follow-up.
- ai-memory: only commands/hook*.rs and commands/install_hooks.rs are
  unchanged in the v2.3.2...v2.4.0 compare (config.rs, serve.rs, run.rs,
  reindex.rs, mcp_bridge.rs and six ai-memory-core files changed).
- ai-memory v2.4.1 (2026-09-25T21:46:32Z) carries the #792 fix and the
  forward-only V67 migration but not #859; it meets the "a 2.4.1
  appears" condition, and qualifying it (another rehearsal and the
  user's approval) is queued. Recorded in both documents.
- The embedder section cites issue #274 (the Mac-only harness, A17, "the
  workstation now owns R1") and catalogs/foundation/memory-stack-20260925.json;
  #274 covers the C3/C4 rerun, and the Nemotron-prefix evaluation comes
  after it.
- The release review gains a dated one-line pointer under its title.
- The rtk and markitdown receipts are linked now that #291 has merged.

Final round (2026-09-26):
- ai-memory: the record and the release review state that 2.4.1 was
  rehearsed and, with the user's approval, switched on 2026-09-26
  (00:14:12Z stop, 00:19:51Z start, V66 to V67), replacing the "needs
  another rehearsal and the user's approval" sentences and "2.4.0 serves".
- headroom: 0.39.0 was qualified and is not switched. The reasons are
  source-checked with gh at 00:28Z: the omission notice counts levels
  over all lines (log_compressor.py:461-487 and the Rust format_output;
  same Python blob since v0.37.0; #3635 changed only the Rust file), the
  Rust detector lacks the #3736 guard, and #299/#296 run 0.37.0.
- The embedder section cites experiment.json and convergence.json for
  the C1 to C3 arms (the catalog has no C1 label).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
…acts and receipt)

Qualified on nativestack-5975wx-20260925 against the pinned 0.37.0: the
verified wheel in a new prefix (RECORD 555/555), the unchanged upstream
#3736 test 6/6 (0.37.0: 4/6), and the stack's MCP path over stdio with
byte-exact retrieval and cross-version store reads; an independent
verifier agreed. The coordinator retains 0.37.0: at v0.39.0 the
'[N lines omitted: ...]' notice counts levels over all input lines
(log_compressor.py:461-487 and the Rust format_output; checked with gh),
so fixture A's '[593 lines omitted: 1 ERROR]' reports the kept CRITICAL
line as omitted; the Rust detector still lacks the #3736 timestamp
guard; #299 pins 0.37.0 in the bootstrap and #296's counters run at it.
No pin changes.

- evidence/artifacts/sota-refresh-20260925/headroom-039/: harness and
  verifier scripts, key results as files, all other outputs and the
  production snapshots bundled, a read-only read-back and results.json
  (decision reasons, overturn condition, beacon note, corrections).
- evidence/receipts/headroom-039-qualification-20260925.json
  (native_cli_e2e).

manifests/evidence.json registration follows in the branch's last commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
…station SOTA refresh

docs/decisions/2026-09-25-ai-memory-2-4-0-release-review.md gains a dated
correction section; the original text is kept. GHSA-9pj6-vhgr-3mwh is not
reachable on this host (stateless transport: GET /mcp 405 with allow:
POST, no --http-stateful), and 2.4.0 writes no automatic pre-migration
archive for a store created by 2.3.2 (snapshot_before_db_migration is
gated to the 1.x -> 2.0 upgrade). It also records that the upgrade was
performed on 2026-09-25 and cites the new receipt.

docs/decisions/2026-09-25-workstation-sota-refresh.md is the dated
research-convergence record for the whole workstation-lane refresh: per
unit, the decision, its primary sources (releases, compares, advisories
and the refuter's key finding), the alternatives and the overturn
condition. rtk, markitdown, mcporter and ai-memory were qualified and
switched; socraticode 1.15.0 is qualified with the cutover deferred to no
earlier than 2026-10-01; dagu 2.17.2 is held by the trading lane;
headroom stays 0.37.0 because 0.38.0 carries #3736 on the MCP compress
path; qmd, qdrant, ccusage, worktrunk and vllm are already current;
llama.cpp b11146 is recorded, not qualified; both models are retained;
the ai-memory Nemotron embedder is evaluation-only. The research and
qualification results are cited as private inputs only.

Review repairs (2026-09-25, independent review needs_changes), each
fact re-checked with gh api at about 22:35Z:
- headroom: v0.39.0 (2026-09-25T19:29:55Z) contains d971f7c3 (compare
  ahead 86, behind 0; the release notes list #3748), so the overturn
  condition is met; qualifying 0.39.0 is the queued follow-up.
- ai-memory: only commands/hook*.rs and commands/install_hooks.rs are
  unchanged in the v2.3.2...v2.4.0 compare (config.rs, serve.rs, run.rs,
  reindex.rs, mcp_bridge.rs and six ai-memory-core files changed).
- ai-memory v2.4.1 (2026-09-25T21:46:32Z) carries the #792 fix and the
  forward-only V67 migration but not #859; it meets the "a 2.4.1
  appears" condition, and qualifying it (another rehearsal and the
  user's approval) is queued. Recorded in both documents.
- The embedder section cites issue #274 (the Mac-only harness, A17, "the
  workstation now owns R1") and catalogs/foundation/memory-stack-20260925.json;
  #274 covers the C3/C4 rerun, and the Nemotron-prefix evaluation comes
  after it.
- The release review gains a dated one-line pointer under its title.
- The rtk and markitdown receipts are linked now that #291 has merged.

Final round (2026-09-26):
- ai-memory: the record and the release review state that 2.4.1 was
  rehearsed and, with the user's approval, switched on 2026-09-26
  (00:14:12Z stop, 00:19:51Z start, V66 to V67), replacing the "needs
  another rehearsal and the user's approval" sentences and "2.4.0 serves".
- headroom: 0.39.0 was qualified and is not switched. The reasons are
  source-checked with gh at 00:28Z: the omission notice counts levels
  over all lines (log_compressor.py:461-487 and the Rust format_output;
  same Python blob since v0.37.0; #3635 changed only the Rust file), the
  Rust detector lacks the #3736 guard, and #299/#296 run 0.37.0.
- The embedder section cites experiment.json and convergence.json for
  the C1 to C3 arms (the catalog has no C1 label).

Delta review (2026-09-26): the refresh record says 2.4.1 arm 28/28
(results/step4-rehearsal.json has 28 checks), and the release review's
"now name 2.4.0" is dated to the 2.4.0 cutover and marked superseded by
2.4.1.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
…acts and receipt)

Qualified on nativestack-5975wx-20260925 against the pinned 0.37.0: the
verified wheel in a new prefix (RECORD 555/555), the unchanged upstream
#3736 test 6/6 (0.37.0: 4/6), and the stack's MCP path over stdio with
byte-exact retrieval and cross-version store reads; an independent
verifier agreed. The coordinator retains 0.37.0: at v0.39.0 the
'[N lines omitted: ...]' notice counts levels over all input lines
(log_compressor.py:461-487 and the Rust format_output; checked with gh),
so fixture A's '[593 lines omitted: 1 ERROR]' reports the kept CRITICAL
line as omitted; the Rust detector still lacks the #3736 timestamp
guard; #299 pins 0.37.0 in the bootstrap and #296's counters run at it.
No pin changes.

- evidence/artifacts/sota-refresh-20260925/headroom-039/: harness and
  verifier scripts, key results as files, all other outputs and the
  production snapshots bundled, a read-only read-back and results.json
  (decision reasons, overturn condition, beacon note, corrections).
- evidence/receipts/headroom-039-qualification-20260925.json
  (native_cli_e2e).

manifests/evidence.json registration follows in the branch's last commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 26, 2026
…station SOTA refresh

docs/decisions/2026-09-25-ai-memory-2-4-0-release-review.md gains a dated
correction section; the original text is kept. GHSA-9pj6-vhgr-3mwh is not
reachable on this host (stateless transport: GET /mcp 405 with allow:
POST, no --http-stateful), and 2.4.0 writes no automatic pre-migration
archive for a store created by 2.3.2 (snapshot_before_db_migration is
gated to the 1.x -> 2.0 upgrade). It also records that the upgrade was
performed on 2026-09-25 and cites the new receipt.

docs/decisions/2026-09-25-workstation-sota-refresh.md is the dated
research-convergence record for the whole workstation-lane refresh: per
unit, the decision, its primary sources (releases, compares, advisories
and the refuter's key finding), the alternatives and the overturn
condition. rtk, markitdown, mcporter and ai-memory were qualified and
switched; socraticode 1.15.0 is qualified with the cutover deferred to no
earlier than 2026-10-01; dagu 2.17.2 is held by the trading lane;
headroom stays 0.37.0 because 0.38.0 carries #3736 on the MCP compress
path; qmd, qdrant, ccusage, worktrunk and vllm are already current;
llama.cpp b11146 is recorded, not qualified; both models are retained;
the ai-memory Nemotron embedder is evaluation-only. The research and
qualification results are cited as private inputs only.

Review repairs (2026-09-25, independent review needs_changes), each
fact re-checked with gh api at about 22:35Z:
- headroom: v0.39.0 (2026-09-25T19:29:55Z) contains d971f7c3 (compare
  ahead 86, behind 0; the release notes list #3748), so the overturn
  condition is met; qualifying 0.39.0 is the queued follow-up.
- ai-memory: only commands/hook*.rs and commands/install_hooks.rs are
  unchanged in the v2.3.2...v2.4.0 compare (config.rs, serve.rs, run.rs,
  reindex.rs, mcp_bridge.rs and six ai-memory-core files changed).
- ai-memory v2.4.1 (2026-09-25T21:46:32Z) carries the #792 fix and the
  forward-only V67 migration but not #859; it meets the "a 2.4.1
  appears" condition, and qualifying it (another rehearsal and the
  user's approval) is queued. Recorded in both documents.
- The embedder section cites issue #274 (the Mac-only harness, A17, "the
  workstation now owns R1") and catalogs/foundation/memory-stack-20260925.json;
  #274 covers the C3/C4 rerun, and the Nemotron-prefix evaluation comes
  after it.
- The release review gains a dated one-line pointer under its title.
- The rtk and markitdown receipts are linked now that #291 has merged.

Final round (2026-09-26):
- ai-memory: the record and the release review state that 2.4.1 was
  rehearsed and, with the user's approval, switched on 2026-09-26
  (00:14:12Z stop, 00:19:51Z start, V66 to V67), replacing the "needs
  another rehearsal and the user's approval" sentences and "2.4.0 serves".
- headroom: 0.39.0 was qualified and is not switched. The reasons are
  source-checked with gh at 00:28Z: the omission notice counts levels
  over all lines (log_compressor.py:461-487 and the Rust format_output;
  same Python blob since v0.37.0; #3635 changed only the Rust file), the
  Rust detector lacks the #3736 guard, and #299/#296 run 0.37.0.
- The embedder section cites experiment.json and convergence.json for
  the C1 to C3 arms (the catalog has no C1 label).

Delta review (2026-09-26): the refresh record says 2.4.1 arm 28/28
(results/step4-rehearsal.json has 28 checks), and the release review's
"now name 2.4.0" is dated to the 2.4.0 cutover and marked superseded by
2.4.1.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 26, 2026
…fresh record (headroom 0.39.0 qualified, not switched) (#307)

* ai-memory 2.4.0 qualification evidence (sanitized artifacts, cutover facts and receipt)

Qualified on nativestack-5975wx-20260925 against the installed 2.3.2 on
restored copies of an online backup (2.3.2 baseline 16/16 on its second
attempt, 2.4.0 rehearsal 19/19: V64 -> V66 with nothing missing, /healthz
200, 23 tools unchanged, both hook clients captured, 2.3.2 refuses the
migrated store), reviewed by an independent verifier (agree, 3 minor
defects), then switched by the coordinator in a user-approved cold-copy
cutover (20:43:48Z stop, 20:44:37Z start, V65/V66 applied, fingerprint
PASS, real Claude Code capture). The Codex /hooks trust step for the 7
changed commands is pending and the user's.

- evidence/artifacts/sota-refresh-20260925/ai-memory/: the driver and
  the verifier's scripts byte-identical, sanitized step results,
  cutover.json with the coordinator's aggregate facts only (no cold copy,
  fingerprint file or store content), a read-only read-back and
  results.json with both arms, the verifier verdict and what was not
  published.
- evidence/receipts/ai-memory-240-qualification-20260925.json
  (native_cli_e2e).

Review repairs (2026-09-25, independent review needs_changes): the
receipt adds the failed production compare check (claude_settings_unchanged:
~/.claude/settings.json changed at 19:01:59Z by an unidentified writer);
results.json corrects the verifier's two wording defects in place and keeps
the originals in a corrections list; recorded_at_utc is the time the
receipt content was written. Final round: the #792 limitation says the
fix is not in 2.4.0 and was released in v2.4.1 at 2026-09-25T21:46:32Z.

manifests/evidence.json registration follows in the branch's last commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* mcporter 0.14.1 qualification evidence (sanitized artifacts, cutover facts and receipt)

Qualified on nativestack-5975wx-20260925 against the installed 0.13.13:
the verified npm tarball (sha256 8e489864..., SLSA provenance from
release.yml at refs/tags/v0.14.1) in a new prefix, and the same 37-step
acceptance in isolated daemon namespaces passing 37/37 on both versions
with identical outcomes. An independent verifier agreed and measured the
one difference the acceptance cannot see: without ps on the daemon's
PATH, 0.14.1 cannot connect keep-alive servers and `daemon stop`
refuses. The coordinator then relinked bin/mcporter with no daemon
running; `list socraticode --brief --no-oauth` listed 26 tools and
codebase_health reached Qdrant.

- evidence/artifacts/sota-refresh-20260925/mcporter/: harness and
  verifier scripts byte-identical, both arms' results, the per-step raw
  outputs bundled into two JSON files (deduplicated), integrity and
  provenance outputs, cutover.json, a read-only read-back and
  results.json.
- evidence/receipts/mcporter-0141-qualification-20260925.json
  (native_cli_e2e).

Review repairs (2026-09-25, independent review needs_changes): the
staging install's package.json and package-lock.json are published as
stage/package.evidence.json and stage/package-lock.evidence.json (bytes
unchanged), so they are not repository dependency manifests and the OSV
inventory and Dependabot scope stay as they were; results.json records
the renames; recorded_at_utc is the time the receipt content was written.

Delta review (2026-09-26): the two npm registry key .pem files are
listed under not_published (source https://registry.npmjs.org/-/npm/v1/keys),
because .gitignore excludes *.pem; they were never committed.

manifests/evidence.json registration follows in the branch's last commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ai-memory 2.4.1 qualification and cutover evidence (sanitized artifacts and receipt)

Qualified on nativestack-5975wx-20260925 against production 2.4.0 on
restored copies of one online backup (2.4.0 arm 26/26, 2.4.1 arm 28/28:
V66 -> V67 with nothing missing, 2.4.0 refuses the V67 copy, path-only
install-hooks render, and the #792 A/B: 65/65 keepalive timers and all 60
vanished peers reclaimed on 2.4.1, none on 2.4.0), reviewed by an
independent verifier (agree, 8 minor defects). The coordinator then
performed the user-approved cutover on 2026-09-26: stop 00:14:12Z, cold
V66 copy, relink, path-only hook rewrite, start 00:19:51Z with V67
applied, fingerprint PASS, keepalive on 2 of 2 live sockets and real
Claude Code capture. The Codex /hooks re-trust is pending and the
user's. The 2.4.0 receipt stays as history.

- evidence/artifacts/sota-refresh-20260925/ai-memory-241/: the driver
  and the verifier's scripts byte-identical, sanitized results for both,
  cutover.json with the coordinator's aggregate facts only (no cold copy,
  fingerprint file or store content), a read-only read-back and
  results.json (steps for both arms, the verdict, and the verifier's
  wording corrections).
- evidence/receipts/ai-memory-241-qualification-20260925.json
  (native_cli_e2e); recorded_at_utc is the time the content was written.

Delta review (2026-09-26): results/step4-rehearsal.json has 28 checks,
all true, so the receipt and the B-arm excerpt say 28 (the original 29 is
kept in results.json corrections), and the R1/A17 control-build limitation
cites issue #274.

manifests/evidence.json registration follows in the branch's last commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* headroom 0.39.0 qualification evidence, not switched (sanitized artifacts and receipt)

Qualified on nativestack-5975wx-20260925 against the pinned 0.37.0: the
verified wheel in a new prefix (RECORD 555/555), the unchanged upstream
#3736 test 6/6 (0.37.0: 4/6), and the stack's MCP path over stdio with
byte-exact retrieval and cross-version store reads; an independent
verifier agreed. The coordinator retains 0.37.0: at v0.39.0 the
'[N lines omitted: ...]' notice counts levels over all input lines
(log_compressor.py:461-487 and the Rust format_output; checked with gh),
so fixture A's '[593 lines omitted: 1 ERROR]' reports the kept CRITICAL
line as omitted; the Rust detector still lacks the #3736 timestamp
guard; #299 pins 0.37.0 in the bootstrap and #296's counters run at it.
No pin changes.

- evidence/artifacts/sota-refresh-20260925/headroom-039/: harness and
  verifier scripts, key results as files, all other outputs and the
  production snapshots bundled, a read-only read-back and results.json
  (decision reasons, overturn condition, beacon note, corrections).
- evidence/receipts/headroom-039-qualification-20260925.json
  (native_cli_e2e).

manifests/evidence.json registration follows in the branch's last commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Linux pins, hook template, recipes and docs: ai-memory 2.4.1 and mcporter 0.14.1

adoption/pins-linux-x86_64.json moves ai-memory to the v2.4.1 archive
(sha256 15cafdc4..., 16,026,020 bytes, equal to the release sidecar, the
API digest and the release body; binary dec065a8...; tag commit 433a19f3) and mcporter to the 0.14.1 npm
tarball (sha256 8e489864..., integrity sha512-/6pv...). The notes carry
the forward-only migration and the ps-on-PATH requirement.
pins-macos-arm64.json is unchanged: the Mac qualifies on its own.

adoption/templates/claude.settings.template.json: the eight ai-memory
hook commands name tools/ai-memory-2.4.1, matching the live source host
after the upstream install-hooks rewrite. bootstrap.md,
platforms/linux-wsl2.md and platforms/macos-arm64.md say the Linux pin
file and the template changed after v2026.09.25.2, and tell a Mac to
point the rendered hooks at its installed 2.3.2 prefix.

recipes/README.md: both rows name the new pins, and a new "Upgrading an
existing store" section gives the cold-copy procedure the host used.
Copies follow in docs/token-efficiency-stack.json,
tools/token-report/README.md and observability/native-data/README.md.
The two catalogs/landscape/upstream-snapshot.json rows come from a
fresh run of their gh api endpoints, and the saturation-audit rows cite
the new receipts. catalogs/landscape/foundation.json winner pins do not
move.

tests/test_adoption_bootstrap_macos.py required every shared component to
carry the same version and npm hash on both platforms, which a Linux-only
move breaks. It now names each Mac-behind-Linux lag exactly (ai-memory
2.3.2/2.4.0 and mcporter 0.13.13/0.14.1, with their receipts), so a move
on either side fails until the table is reviewed, and compares npm hashes
only between equal versions.

Review repairs (2026-09-25, independent review needs_changes):
- recipes/README.md: the hook commands' 49374 is an example; use the
  host's own server URL (NativeStack binds 127.0.0.1:49474, and 49374 is
  another distro's default there) and record what NativeStack used. The
  upgrade section states the order hazard: bootstrap-linux.sh repoints
  bin/ai-memory at once, so an existing store is stopped and cold-copied
  before any bootstrap re-run; adoption/update.md step 1 says the same.
- platforms/macos-arm64.md: the changed-after marker in step 3 states only
  the difference and points to a standing "ai-memory hook paths on macOS"
  note, which gives the full recipe command (--apply, --capture-mode
  allowlist, --no-capture-prompts, --server-url): with no stored mode,
  v2.3.2's installer falls back to denylist.
- catalogs/landscape/foundation.json durable-memory: one appended, dated
  limitation for the 20:44Z cutover; the earlier 2026-09-25 text is kept,
  and the winner pin (a verdict field) does not change.

Final round (2026-09-26): the ai-memory pin, the hook template, the
recipes and docs, the upstream-snapshot row, the saturation-audit row and
the Mac-lag test pair move from 2.4.0 to 2.4.1 after the user-approved
2.4.1 cutover (the 2.4.0 receipt stays as history). The upgrade section
covers both cutovers and warns against `pgrep -f` in the quiesce wait.
platforms/macos-arm64.md now says accurately that #299 also added
ccusage, headroom, repomix, serena, socraticode and toon to the Linux
pins, and which of the ten have Mac entries. The foundation.json
limitation names both cutovers.

Delta review (2026-09-26): the ai-memory install_note dates the
three-value check to 2026-09-25 (22:43Z qualification, 23:10Z verifier)
and says 2026-09-26 re-checked the API digest.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Decision records: correct the ai-memory 2.4.0 review; record the workstation SOTA refresh

docs/decisions/2026-09-25-ai-memory-2-4-0-release-review.md gains a dated
correction section; the original text is kept. GHSA-9pj6-vhgr-3mwh is not
reachable on this host (stateless transport: GET /mcp 405 with allow:
POST, no --http-stateful), and 2.4.0 writes no automatic pre-migration
archive for a store created by 2.3.2 (snapshot_before_db_migration is
gated to the 1.x -> 2.0 upgrade). It also records that the upgrade was
performed on 2026-09-25 and cites the new receipt.

docs/decisions/2026-09-25-workstation-sota-refresh.md is the dated
research-convergence record for the whole workstation-lane refresh: per
unit, the decision, its primary sources (releases, compares, advisories
and the refuter's key finding), the alternatives and the overturn
condition. rtk, markitdown, mcporter and ai-memory were qualified and
switched; socraticode 1.15.0 is qualified with the cutover deferred to no
earlier than 2026-10-01; dagu 2.17.2 is held by the trading lane;
headroom stays 0.37.0 because 0.38.0 carries #3736 on the MCP compress
path; qmd, qdrant, ccusage, worktrunk and vllm are already current;
llama.cpp b11146 is recorded, not qualified; both models are retained;
the ai-memory Nemotron embedder is evaluation-only. The research and
qualification results are cited as private inputs only.

Review repairs (2026-09-25, independent review needs_changes), each
fact re-checked with gh api at about 22:35Z:
- headroom: v0.39.0 (2026-09-25T19:29:55Z) contains d971f7c3 (compare
  ahead 86, behind 0; the release notes list #3748), so the overturn
  condition is met; qualifying 0.39.0 is the queued follow-up.
- ai-memory: only commands/hook*.rs and commands/install_hooks.rs are
  unchanged in the v2.3.2...v2.4.0 compare (config.rs, serve.rs, run.rs,
  reindex.rs, mcp_bridge.rs and six ai-memory-core files changed).
- ai-memory v2.4.1 (2026-09-25T21:46:32Z) carries the #792 fix and the
  forward-only V67 migration but not #859; it meets the "a 2.4.1
  appears" condition, and qualifying it (another rehearsal and the
  user's approval) is queued. Recorded in both documents.
- The embedder section cites issue #274 (the Mac-only harness, A17, "the
  workstation now owns R1") and catalogs/foundation/memory-stack-20260925.json;
  #274 covers the C3/C4 rerun, and the Nemotron-prefix evaluation comes
  after it.
- The release review gains a dated one-line pointer under its title.
- The rtk and markitdown receipts are linked now that #291 has merged.

Final round (2026-09-26):
- ai-memory: the record and the release review state that 2.4.1 was
  rehearsed and, with the user's approval, switched on 2026-09-26
  (00:14:12Z stop, 00:19:51Z start, V66 to V67), replacing the "needs
  another rehearsal and the user's approval" sentences and "2.4.0 serves".
- headroom: 0.39.0 was qualified and is not switched. The reasons are
  source-checked with gh at 00:28Z: the omission notice counts levels
  over all lines (log_compressor.py:461-487 and the Rust format_output;
  same Python blob since v0.37.0; #3635 changed only the Rust file), the
  Rust detector lacks the #3736 guard, and #299/#296 run 0.37.0.
- The embedder section cites experiment.json and convergence.json for
  the C1 to C3 arms (the catalog has no C1 label).

Delta review (2026-09-26): the refresh record says 2.4.1 arm 28/28
(results/step4-rehearsal.json has 28 checks), and the release review's
"now name 2.4.0" is dated to the 2.4.0 cutover and marked superseded by
2.4.1.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Stack pins ai-memory 2.4.1 and mcporter 0.14.1; register evidence and regenerate reports

Rebuilt on main c09dd6d (#306) per the hot-file protocol: the reviewed
stack.json diff, the same registered path set and all four qualification
receipts as the reviewed head b64f9b0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 26, 2026
… vLLM 0.30.0 embedder switch

Mirrors the workstation's #296 method on wsl-authoring-20260923 (13 of #296's 16 tasks; 29
children, 15 Sonnet 5 + 14 Opus 5.5 at effort max). Publishes each tool's check result and
independent verifier verdict, including the refuted rtk claim (branch -a '+' marker; the
bare git show exclusion missing the -C form), Serena's missed cross-file references and
repomix --compress dropping a def. Adds a baseline_kind field so only counterfactual_raw
pairs read as reductions. Adds the laptop's vLLM 0.25.0 -> 0.30.0 qualification and switch
receipt, reporting set identity, strict-order failures and the production-vs-production
control totals.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 26, 2026
… vLLM 0.30.0 embedder switch (#316)

Mirrors the workstation's #296 method on wsl-authoring-20260923 (13 of #296's 16 tasks; 29
children, 15 Sonnet 5 + 14 Opus 5.5 at effort max). Publishes each tool's check result and
independent verifier verdict, including the refuted rtk claim (branch -a '+' marker; the
bare git show exclusion missing the -C form), Serena's missed cross-file references and
repomix --compress dropping a def. Adds a baseline_kind field so only counterfactual_raw
pairs read as reductions. Adds the laptop's vLLM 0.25.0 -> 0.30.0 qualification and switch
receipt, reporting set identity, strict-order failures and the production-vs-production
control totals.

Co-authored-by: seathatflowsinourveins <234074349+seathatflowsinourveins@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 26, 2026
…ecks, four real limits, and a correction to #296's QMD row (#343)

* Record the token stack inside native Codex sessions, and correct #296's QMD row

Codex side of #296: one native `codex exec` session per tool (codex-cli
0.155.1, gpt-6-astra, effort max) in a dedicated worktree, using RTK's
standing instructions, the Context Mode plugin, the #289 user-scope MCP
servers and the CLIs. 11 of 15 tools pass their answer checks; RTK +17782
saved for the run-only worktree; nine exact o200k comparisons from 28.2% to
99.8%; Codex-returned usage 7,085,015 input (6,253,312 cached) and 164,906
output over 17 sessions. The four failures are real limits: jCodeMunch is
project-scoped (#240), QMD's catalog index excludes adoption/, Repomix
--compress dropped a declaration, Context Hub lacks the upstream facts.
#296's QMD PASS retrieved a wrong document for its question; its README now
says so.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Sanitize the Codex E2E receipt after cross-family review (PR #343)

Codex review (gpt-6-astra, read-only) confirmed every number and evidence
class and asked for three fixes: replace the live-clone path, the
session-derived codebase-memory project id and the SocratiCode collection id
with placeholders, and state the QMD document count per attempt (121, then
122 as the live index grew).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 27, 2026
…asks,

metrics, gate, procedure)

Freezes the E2E preregistration half of full-save/PLAN.md step 9 (the
baseline-receipt half already merged via #369) before any organic run:
77 frozen task definitions (16 reused Claude + 15 reused Codex information
needs from #296/#343's receipts, 46 seeded/control tasks), every AA §8.3
metric/threshold/guardrail/outcome rule with its three "proposed" bounds now
stated as frozen, the AA §8.1b capability gate, the AA §8.4 procedure, the
AA §8.5 overturn conditions, the identity scheme, the Q1 blind-process
qualification, the Workflow script that will run arms B/A/A0 (syntax-checked
only, never executed), and the measurement runbook.

Corrected on independent Claude review (11 findings, 95 turns) and a GPT-6
cross-family review (2 findings, 1 high) plus one combined repair round (13/13
addressed with failing-first evidence): relocated from
blueprints/token-practice-e2e/ to evidence/artifacts/token-adoption-e2e-20260926/
per adoption-audit/PLAN.md line 402; moved the Workflow script out of
examples/claude-native/workflows/ to avoid that directory's shared
PACKET/ROUTING contract test, which this one-off frozen script correctly does
not match; corrected PR-A's status (its base branch merged as #369 -- baseline
only, tooling-fix scope still open); added the six AA §6 machine-readable
policy fields the JSON was missing; fixed a wrong Codex arm-A lifecycle
binding; blocked the one information need (an otel-tui receiver task) that no
current role may legally run, pending a named permitted role; froze the
coordinator's own main-dispatch task at xhigh instead of max; fixed a
denylist word-boundary gap that let underscore-qualified MCP names slip past;
and sealed the five normative files' SHA256 with an append-only amendment
rule. Also refreshes PR #376 and #364's status to merged (both landed while
this unit was in review), while recording that #364's field-preservation gap
(workflow.run_id/tool_use_id still missing from collector.yaml) persists after
its merge.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 27, 2026
…asks, (#381)

metrics, gate, procedure)

Freezes the E2E preregistration half of full-save/PLAN.md step 9 (the
baseline-receipt half already merged via #369) before any organic run:
77 frozen task definitions (16 reused Claude + 15 reused Codex information
needs from #296/#343's receipts, 46 seeded/control tasks), every AA §8.3
metric/threshold/guardrail/outcome rule with its three "proposed" bounds now
stated as frozen, the AA §8.1b capability gate, the AA §8.4 procedure, the
AA §8.5 overturn conditions, the identity scheme, the Q1 blind-process
qualification, the Workflow script that will run arms B/A/A0 (syntax-checked
only, never executed), and the measurement runbook.

Corrected on independent Claude review (11 findings, 95 turns) and a GPT-6
cross-family review (2 findings, 1 high) plus one combined repair round (13/13
addressed with failing-first evidence): relocated from
blueprints/token-practice-e2e/ to evidence/artifacts/token-adoption-e2e-20260926/
per adoption-audit/PLAN.md line 402; moved the Workflow script out of
examples/claude-native/workflows/ to avoid that directory's shared
PACKET/ROUTING contract test, which this one-off frozen script correctly does
not match; corrected PR-A's status (its base branch merged as #369 -- baseline
only, tooling-fix scope still open); added the six AA §6 machine-readable
policy fields the JSON was missing; fixed a wrong Codex arm-A lifecycle
binding; blocked the one information need (an otel-tui receiver task) that no
current role may legally run, pending a named permitted role; froze the
coordinator's own main-dispatch task at xhigh instead of max; fixed a
denylist word-boundary gap that let underscore-qualified MCP names slip past;
and sealed the five normative files' SHA256 with an append-only amendment
rule. Also refreshes PR #376 and #364's status to merged (both landed while
this unit was in review), while recording that #364's field-preservation gap
(workflow.run_id/tool_use_id still missing from collector.yaml) persists after
its merge.

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 27, 2026
…ta from the 2026-09-27 verdict wave (#401)

* E1/E2 token E2E receipts: dated adjudications, errata and reproduced answer oracles

Applies the 2026-09-27 cross-family GPT-6 review of the E1 (#296) and E2 (#316)
receipts, append-only: every recorded field keeps its value, and each receipt keeps
its own serialization.

- Every tools[] row gains adjudication {as_of, task_acceptance, bad_input_control_recorded,
  basis, evidence}, using the review's status table and vocabulary (partial,
  retracted_later, untested, unsupported).
- E1: QMD retracted_later (wrong_document, corrections_to_296); Repomix retracted_later
  (count=47 PASS on a 48-function file; repomix v1.18.1 PythonParseStrategy.ts L81-86
  tests only a def's first line); ai-memory untested (base_empty=True is a vacuous pass);
  the other thirteen partial with no recorded failing control. A dated errata block
  names the pre-errata sha256 the preregistration records as a historical identity.
- E2: context-mode, jcodemunch, qmd and ast-grep stay partial with no recorded failing
  control; SocratiCode's comparison counts an abridged transcription; RTK retracted by
  its verifier; Repomix's FAIL accurate.
- READMEs carry dated corrections; the review is retained, sanitized, beside E1.
- tests/test_token_e2e_receipt_checks.py: receipt contracts (red before the errata) and
  answer oracles rebuilt from pinned Git blobs, hash-matched to the recorded baselines,
  each with failing controls (red under a weaker historical-style check).

Sources: docs/acceptance-evidence-policy.md (Discriminating controls);
https://github.com/yamadashy/repomix/blob/v1.18.1/src/core/treeSitter/parseStrategies/PythonParseStrategy.ts#L81
evidence/artifacts/token-e2e-codex-20260926/receipt.json#/corrections_to_296.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Token layer: routing rules and recipe limits for known upstream defects

Closes the documentation, recipe and routing gaps from the 2026-09-27 per-tool
verdict wave for the token-efficiency layer. Every rule cites the pinned upstream
source or a committed receipt; no pin changes.

recipes/README.md (token-tool rows and sections):
- agentsview: v0.44.0 unqualified; worker populations need --include-children,
  --include-automated and --include-one-shot, and --fts reads message bodies only
  (v0.43.0 docs/session-api.md, docs/commands.md); archive answers are observation.
- ast-grep: every 0.45.3 command registers customLanguages from any sgconfig.yml in
  the working directory or a parent (lib.rs L107-139, config.rs L98-131/L280-299);
  PR #2960's opt-in is unreleased; outline (#2957) and YAML rule (#2963) limits.
- ccusage: token-only while any model is unpriced (v20.0.24 json-output.md
  Unpriced Models; config-files.md pricingOverrides).
- codebase-memory: trace_path excludes tests and evidence by default, search_code
  hits can be mentions (v0.11.0 src/mcp/mcp.c schemas); caller-set procedure.
- context-hub: check metadata.versions (FastAPI guide 0.136.3 vs 0.141.1).
- context-mode: npm tarball digest and plugin commits are separate identities.
- headroom: v0.39.1 changes only the proxy limiter, so the omission-count defect stands.
- markitdown: HTML as .txt passes through; -x html / -m text/html; selective extras.
- mcporter: --name does not select the server; a positional token becomes the
  selector (v0.14.1 call-arguments.ts L164-192, ephemeral-flags.ts L97-105); the
  jCodeMunch example gains --server.
- qmd: MCP query expands and reranks by default (v2.8.3 README, server.ts); typed
  lex + rerank:false + bounded get; catalog scope; refresh guard for #989/#991
  (src/store.ts reindexCollection).
- repomix: multi-line Python signatures are dropped (v1.18.1 PythonParseStrategy.ts
  L81-86); exact-definition tasks go to an uncompressed pack or the original.
- toon: BOM root-string round trip (#339); no record-count threshold upstream;
  nested-uniform columns are tabular (tabular.ts L61-78).

docs/token-session-handbook.md: a "Known upstream limits behind the lanes" section
outside the carrier text (every carrier line unchanged), and row pointers.
docs/token-practice.md: counts, comparisons and acceptance rules (invocation counts
are not success rates; comparisons need task acceptance; count recovery reads;
gpt-tokenizer's special-token contract; receipts carry adjudications).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Apply the GPT-6 review of this branch: exact inventory oracle, TOON eligibility wording

One cross-family review round (GPT-6 through codex-omniroute, read-only) returned
DEFECTS FOUND with two findings; both are fixed:

- P2 tests/test_token_e2e_receipt_checks.py: inventory_oracle checked only the claimed
  count plus subset membership, so a correct count with missing, duplicated or no names
  passed. It now requires the exact name set with one entry per function. New controls keep
  the correct count while dropping, duplicating or inventing a name, or listing none; they
  failed against the old oracle first (red) and pass now.
- P3 docs/token-session-handbook.md (and the matching recipes/README.md toon row): "accepts
  any non-empty uniform array" overstated TOON 4.1.1's tabular eligibility. The encoder
  needs non-empty objects sharing one key set, with columns of primitives or, recursively,
  of non-empty objects sharing one key set; an empty object or an array-valued column falls
  back to list form (packages/toon/src/encode/tabular.ts L6-78 at v4.1.1; confirmed with the
  installed 4.1.1 CLI on [{}], [{"a":[]}], [{"a":{}}] and [{"a":{"b":1}}]).

Also: the ast-grep oracle keys counts by path under scripts/ rather than file stem, and E1's
Context Mode adjudication states that attempt 1's marker-grep FAIL on a refused call is not
shown to be a run of attempt 2's answer check (the receipt does not record that command).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Apply the independent verifier's findings: TOON source line, citation anchors, E2 Serena erratum

The independent verifier of this branch reported five defects; all are fixed:

- docs/token-session-handbook.md, the TOON Sources line under the carrier (on main
  before this branch) still said "The encoder accepts a non-empty uniform array",
  the wording GPT-6 finding P3 refuted. It now states the encoder's rule from
  packages/toon/src/encode/tabular.ts L6-78 at v4.1.1 and links the handbook's
  "Known upstream limits behind the lanes" section. It is not a "- " carrier line,
  so tests/test_token_lanes_subagent_start.py's verbatim-line check is unaffected.
- Citation anchors: SocratiCode's INCLUDE_DOT_FILES default is documented under
  README.md "Indexing Behaviour" (v1.14.0 L1574-1579), not "Ignore Rules", so the
  handbook links #indexing-behaviour. codebase-memory-mcp v0.11.0 search_code gets
  its own link to src/mcp/mcp.c#L632-L659 in the handbook, and the recipe gives
  index_repository (L466-481), trace_path (L532-560) and search_code (L632-659)
  one link each. Both upstream files matched the retained copies by sha256 on
  2026-09-27.
- E2 receipt errata.items[4] and the E2 README erratum now name Serena's accurate
  reference FAIL (find_referencing_symbols missed 6 of 8 call sites; the definition
  check passed) and give each group of partial rows its own reason. Only that one
  finding string changed in receipt.json (same serialization, 1 line); written_at
  is unchanged because this refines the unmerged 2026-09-27 erratum on the same day.
- tests/test_token_e2e_receipt_checks.py: the require_commit docstring now says the
  guard is modelled on RetainedEvidenceTests' Git check and adds the pinned-commit
  probe (tests/test_adoption_status.py L1721-1724 checks only git and the checkout).

Sources: toon-format/toon v4.1.1 packages/toon/src/encode/tabular.ts;
giancarloerra/SocratiCode v1.14.0 README.md; DeusData/codebase-memory-mcp v0.11.0
src/mcp/mcp.c; the E2 receipt's own serena row (tools[5]) and README.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant