Repository navigation
Token stack inside Ultracode subagents: 16/16 tools E2E with lifetime counters, exact comparisons and five new host receipts - #296
Merged
seathatflowsinourveins merged 3 commits intoSep 25, 2026
Conversation
…-5975wx-20260925 Install and use-stage receipts for ast-grep 0.45.3, codebase-memory-mcp 0.11.0, context-hub 0.1.4, agentsview 0.43.0, and otel-tui 0.7.5, each pinned to their manifests/stack.json versions and recorded with scripts/host_receipts.py per docs/contributing-evidence.md. Each use-stage command asserts its own result in-line (a positive control plus a cheap negative control) so the printed output backs the claim without relying on a human to eyeball a clean run. The codebase-memory-mcp use receipt was superseded once to keep its assertion lines inside the 400-char excerpt window; the original stays byte-identical. Refreshes the derived component-evidence-matrix and evidence manifest. All receipts carry only the recorder's self review; independent review is still required before any of these count toward accepted status. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ime counters Sixteen Sonnet 5 workflow subagents each used one selected tool on real work in this repository and checked the answer against the plain baseline (all 16 pass; Context Mode, Serena and Context Hub on a second attempt after a real binding constraint). Upstream counters moved over the 28-minute window: RTK +56,527 saved for this run's worktree alone and Headroom +183,904, which the exact o200k comparison of the same compression matches at 180,752. Thirteen exact per-task comparisons, per-child provider usage, a fresh-session MCP check and the retained failures are in the receipt. Codex workers were not run: the account is at its usage limit until 2026-09-30. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
… (PR #296) A separate headless Claude session (Sonnet) recorded agree on all 11 receipts (ast-grep, codebase-memory-mcp, context-hub, agentsview, otel-tui install/use). It found one mismatch in the E2E README: the fresh-session server list omitted the claude.ai Claude Docs connector, now listed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
deleted the
claude/token-e2e-ultracode-20260925
branch
September 25, 2026 22:40
1 task done
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 25, 2026
… request lane (#298) Two foundation gates. token-stack-subagent-e2e points at #296's receipt: 16 of 16 tools used in Sonnet workflow subagents with passing checks, RTK +56527 for the run-only worktree, Headroom +183904 lifetime, Codex workers blocked until 2026-09-30. host-request-lane points at the #266 recipe: the workstation poll timer runs every 10 min, 0 workstation requests, 2 mac-coordinator requests open (#274, #276). Values use the dashboard token alphabet. Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 25, 2026
… DO_NOT_TRACK=1) (#304) headroom-ai 0.37.0 uploads an anonymous usage beacon by default (telemetry/beacon.py BEACON_DEFAULT_ON = True; ccr/mcp_server.py calls it on every MCP compression). offline.py's HEADROOM_OFFLINE is the master no-egress switch (beacon, update check, usage reporter, HF downloads); DO_NOT_TRACK also disables the beacon. The recipe and token-stack row now set both, with the native Claude registration command; the #296 E2E README records that the beacon was at its default during that run. Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
9 of 22 tasks
Owner
Author
|
Upstream check of two #296 registrations (from the laptop lane, wsl-authoring-20260923; upstream is the source of truth)
The laptop does not register either. Please reconcile the workstation's user-scope set against these upstream statements and the catalog's disposition. |
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
…acts and receipt) Qualified on nativestack-5975wx-20260925 against the pinned 0.37.0: the verified wheel in a new prefix (RECORD 555/555), the unchanged upstream #3736 test 6/6 (0.37.0: 4/6), and the stack's MCP path over stdio with byte-exact retrieval and cross-version store reads; an independent verifier agreed. The coordinator retains 0.37.0: at v0.39.0 the '[N lines omitted: ...]' notice counts levels over all input lines (log_compressor.py:461-487 and the Rust format_output; checked with gh), so fixture A's '[593 lines omitted: 1 ERROR]' reports the kept CRITICAL line as omitted; the Rust detector still lacks the #3736 timestamp guard; #299 pins 0.37.0 in the bootstrap and #296's counters run at it. No pin changes. - evidence/artifacts/sota-refresh-20260925/headroom-039/: harness and verifier scripts, key results as files, all other outputs and the production snapshots bundled, a read-only read-back and results.json (decision reasons, overturn condition, beacon note, corrections). - evidence/receipts/headroom-039-qualification-20260925.json (native_cli_e2e). manifests/evidence.json registration follows in the branch's last commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
…station SOTA refresh docs/decisions/2026-09-25-ai-memory-2-4-0-release-review.md gains a dated correction section; the original text is kept. GHSA-9pj6-vhgr-3mwh is not reachable on this host (stateless transport: GET /mcp 405 with allow: POST, no --http-stateful), and 2.4.0 writes no automatic pre-migration archive for a store created by 2.3.2 (snapshot_before_db_migration is gated to the 1.x -> 2.0 upgrade). It also records that the upgrade was performed on 2026-09-25 and cites the new receipt. docs/decisions/2026-09-25-workstation-sota-refresh.md is the dated research-convergence record for the whole workstation-lane refresh: per unit, the decision, its primary sources (releases, compares, advisories and the refuter's key finding), the alternatives and the overturn condition. rtk, markitdown, mcporter and ai-memory were qualified and switched; socraticode 1.15.0 is qualified with the cutover deferred to no earlier than 2026-10-01; dagu 2.17.2 is held by the trading lane; headroom stays 0.37.0 because 0.38.0 carries #3736 on the MCP compress path; qmd, qdrant, ccusage, worktrunk and vllm are already current; llama.cpp b11146 is recorded, not qualified; both models are retained; the ai-memory Nemotron embedder is evaluation-only. The research and qualification results are cited as private inputs only. Review repairs (2026-09-25, independent review needs_changes), each fact re-checked with gh api at about 22:35Z: - headroom: v0.39.0 (2026-09-25T19:29:55Z) contains d971f7c3 (compare ahead 86, behind 0; the release notes list #3748), so the overturn condition is met; qualifying 0.39.0 is the queued follow-up. - ai-memory: only commands/hook*.rs and commands/install_hooks.rs are unchanged in the v2.3.2...v2.4.0 compare (config.rs, serve.rs, run.rs, reindex.rs, mcp_bridge.rs and six ai-memory-core files changed). - ai-memory v2.4.1 (2026-09-25T21:46:32Z) carries the #792 fix and the forward-only V67 migration but not #859; it meets the "a 2.4.1 appears" condition, and qualifying it (another rehearsal and the user's approval) is queued. Recorded in both documents. - The embedder section cites issue #274 (the Mac-only harness, A17, "the workstation now owns R1") and catalogs/foundation/memory-stack-20260925.json; #274 covers the C3/C4 rerun, and the Nemotron-prefix evaluation comes after it. - The release review gains a dated one-line pointer under its title. - The rtk and markitdown receipts are linked now that #291 has merged. Final round (2026-09-26): - ai-memory: the record and the release review state that 2.4.1 was rehearsed and, with the user's approval, switched on 2026-09-26 (00:14:12Z stop, 00:19:51Z start, V66 to V67), replacing the "needs another rehearsal and the user's approval" sentences and "2.4.0 serves". - headroom: 0.39.0 was qualified and is not switched. The reasons are source-checked with gh at 00:28Z: the omission notice counts levels over all lines (log_compressor.py:461-487 and the Rust format_output; same Python blob since v0.37.0; #3635 changed only the Rust file), the Rust detector lacks the #3736 guard, and #299/#296 run 0.37.0. - The embedder section cites experiment.json and convergence.json for the C1 to C3 arms (the catalog has no C1 label). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
…acts and receipt) Qualified on nativestack-5975wx-20260925 against the pinned 0.37.0: the verified wheel in a new prefix (RECORD 555/555), the unchanged upstream #3736 test 6/6 (0.37.0: 4/6), and the stack's MCP path over stdio with byte-exact retrieval and cross-version store reads; an independent verifier agreed. The coordinator retains 0.37.0: at v0.39.0 the '[N lines omitted: ...]' notice counts levels over all input lines (log_compressor.py:461-487 and the Rust format_output; checked with gh), so fixture A's '[593 lines omitted: 1 ERROR]' reports the kept CRITICAL line as omitted; the Rust detector still lacks the #3736 timestamp guard; #299 pins 0.37.0 in the bootstrap and #296's counters run at it. No pin changes. - evidence/artifacts/sota-refresh-20260925/headroom-039/: harness and verifier scripts, key results as files, all other outputs and the production snapshots bundled, a read-only read-back and results.json (decision reasons, overturn condition, beacon note, corrections). - evidence/receipts/headroom-039-qualification-20260925.json (native_cli_e2e). manifests/evidence.json registration follows in the branch's last commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
…station SOTA refresh docs/decisions/2026-09-25-ai-memory-2-4-0-release-review.md gains a dated correction section; the original text is kept. GHSA-9pj6-vhgr-3mwh is not reachable on this host (stateless transport: GET /mcp 405 with allow: POST, no --http-stateful), and 2.4.0 writes no automatic pre-migration archive for a store created by 2.3.2 (snapshot_before_db_migration is gated to the 1.x -> 2.0 upgrade). It also records that the upgrade was performed on 2026-09-25 and cites the new receipt. docs/decisions/2026-09-25-workstation-sota-refresh.md is the dated research-convergence record for the whole workstation-lane refresh: per unit, the decision, its primary sources (releases, compares, advisories and the refuter's key finding), the alternatives and the overturn condition. rtk, markitdown, mcporter and ai-memory were qualified and switched; socraticode 1.15.0 is qualified with the cutover deferred to no earlier than 2026-10-01; dagu 2.17.2 is held by the trading lane; headroom stays 0.37.0 because 0.38.0 carries #3736 on the MCP compress path; qmd, qdrant, ccusage, worktrunk and vllm are already current; llama.cpp b11146 is recorded, not qualified; both models are retained; the ai-memory Nemotron embedder is evaluation-only. The research and qualification results are cited as private inputs only. Review repairs (2026-09-25, independent review needs_changes), each fact re-checked with gh api at about 22:35Z: - headroom: v0.39.0 (2026-09-25T19:29:55Z) contains d971f7c3 (compare ahead 86, behind 0; the release notes list #3748), so the overturn condition is met; qualifying 0.39.0 is the queued follow-up. - ai-memory: only commands/hook*.rs and commands/install_hooks.rs are unchanged in the v2.3.2...v2.4.0 compare (config.rs, serve.rs, run.rs, reindex.rs, mcp_bridge.rs and six ai-memory-core files changed). - ai-memory v2.4.1 (2026-09-25T21:46:32Z) carries the #792 fix and the forward-only V67 migration but not #859; it meets the "a 2.4.1 appears" condition, and qualifying it (another rehearsal and the user's approval) is queued. Recorded in both documents. - The embedder section cites issue #274 (the Mac-only harness, A17, "the workstation now owns R1") and catalogs/foundation/memory-stack-20260925.json; #274 covers the C3/C4 rerun, and the Nemotron-prefix evaluation comes after it. - The release review gains a dated one-line pointer under its title. - The rtk and markitdown receipts are linked now that #291 has merged. Final round (2026-09-26): - ai-memory: the record and the release review state that 2.4.1 was rehearsed and, with the user's approval, switched on 2026-09-26 (00:14:12Z stop, 00:19:51Z start, V66 to V67), replacing the "needs another rehearsal and the user's approval" sentences and "2.4.0 serves". - headroom: 0.39.0 was qualified and is not switched. The reasons are source-checked with gh at 00:28Z: the omission notice counts levels over all lines (log_compressor.py:461-487 and the Rust format_output; same Python blob since v0.37.0; #3635 changed only the Rust file), the Rust detector lacks the #3736 guard, and #299/#296 run 0.37.0. - The embedder section cites experiment.json and convergence.json for the C1 to C3 arms (the catalog has no C1 label). Delta review (2026-09-26): the refresh record says 2.4.1 arm 28/28 (results/step4-rehearsal.json has 28 checks), and the release review's "now name 2.4.0" is dated to the 2.4.0 cutover and marked superseded by 2.4.1. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
…acts and receipt) Qualified on nativestack-5975wx-20260925 against the pinned 0.37.0: the verified wheel in a new prefix (RECORD 555/555), the unchanged upstream #3736 test 6/6 (0.37.0: 4/6), and the stack's MCP path over stdio with byte-exact retrieval and cross-version store reads; an independent verifier agreed. The coordinator retains 0.37.0: at v0.39.0 the '[N lines omitted: ...]' notice counts levels over all input lines (log_compressor.py:461-487 and the Rust format_output; checked with gh), so fixture A's '[593 lines omitted: 1 ERROR]' reports the kept CRITICAL line as omitted; the Rust detector still lacks the #3736 timestamp guard; #299 pins 0.37.0 in the bootstrap and #296's counters run at it. No pin changes. - evidence/artifacts/sota-refresh-20260925/headroom-039/: harness and verifier scripts, key results as files, all other outputs and the production snapshots bundled, a read-only read-back and results.json (decision reasons, overturn condition, beacon note, corrections). - evidence/receipts/headroom-039-qualification-20260925.json (native_cli_e2e). manifests/evidence.json registration follows in the branch's last commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 26, 2026
…station SOTA refresh docs/decisions/2026-09-25-ai-memory-2-4-0-release-review.md gains a dated correction section; the original text is kept. GHSA-9pj6-vhgr-3mwh is not reachable on this host (stateless transport: GET /mcp 405 with allow: POST, no --http-stateful), and 2.4.0 writes no automatic pre-migration archive for a store created by 2.3.2 (snapshot_before_db_migration is gated to the 1.x -> 2.0 upgrade). It also records that the upgrade was performed on 2026-09-25 and cites the new receipt. docs/decisions/2026-09-25-workstation-sota-refresh.md is the dated research-convergence record for the whole workstation-lane refresh: per unit, the decision, its primary sources (releases, compares, advisories and the refuter's key finding), the alternatives and the overturn condition. rtk, markitdown, mcporter and ai-memory were qualified and switched; socraticode 1.15.0 is qualified with the cutover deferred to no earlier than 2026-10-01; dagu 2.17.2 is held by the trading lane; headroom stays 0.37.0 because 0.38.0 carries #3736 on the MCP compress path; qmd, qdrant, ccusage, worktrunk and vllm are already current; llama.cpp b11146 is recorded, not qualified; both models are retained; the ai-memory Nemotron embedder is evaluation-only. The research and qualification results are cited as private inputs only. Review repairs (2026-09-25, independent review needs_changes), each fact re-checked with gh api at about 22:35Z: - headroom: v0.39.0 (2026-09-25T19:29:55Z) contains d971f7c3 (compare ahead 86, behind 0; the release notes list #3748), so the overturn condition is met; qualifying 0.39.0 is the queued follow-up. - ai-memory: only commands/hook*.rs and commands/install_hooks.rs are unchanged in the v2.3.2...v2.4.0 compare (config.rs, serve.rs, run.rs, reindex.rs, mcp_bridge.rs and six ai-memory-core files changed). - ai-memory v2.4.1 (2026-09-25T21:46:32Z) carries the #792 fix and the forward-only V67 migration but not #859; it meets the "a 2.4.1 appears" condition, and qualifying it (another rehearsal and the user's approval) is queued. Recorded in both documents. - The embedder section cites issue #274 (the Mac-only harness, A17, "the workstation now owns R1") and catalogs/foundation/memory-stack-20260925.json; #274 covers the C3/C4 rerun, and the Nemotron-prefix evaluation comes after it. - The release review gains a dated one-line pointer under its title. - The rtk and markitdown receipts are linked now that #291 has merged. Final round (2026-09-26): - ai-memory: the record and the release review state that 2.4.1 was rehearsed and, with the user's approval, switched on 2026-09-26 (00:14:12Z stop, 00:19:51Z start, V66 to V67), replacing the "needs another rehearsal and the user's approval" sentences and "2.4.0 serves". - headroom: 0.39.0 was qualified and is not switched. The reasons are source-checked with gh at 00:28Z: the omission notice counts levels over all lines (log_compressor.py:461-487 and the Rust format_output; same Python blob since v0.37.0; #3635 changed only the Rust file), the Rust detector lacks the #3736 guard, and #299/#296 run 0.37.0. - The embedder section cites experiment.json and convergence.json for the C1 to C3 arms (the catalog has no C1 label). Delta review (2026-09-26): the refresh record says 2.4.1 arm 28/28 (results/step4-rehearsal.json has 28 checks), and the release review's "now name 2.4.0" is dated to the 2.4.0 cutover and marked superseded by 2.4.1. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 26, 2026
…fresh record (headroom 0.39.0 qualified, not switched) (#307) * ai-memory 2.4.0 qualification evidence (sanitized artifacts, cutover facts and receipt) Qualified on nativestack-5975wx-20260925 against the installed 2.3.2 on restored copies of an online backup (2.3.2 baseline 16/16 on its second attempt, 2.4.0 rehearsal 19/19: V64 -> V66 with nothing missing, /healthz 200, 23 tools unchanged, both hook clients captured, 2.3.2 refuses the migrated store), reviewed by an independent verifier (agree, 3 minor defects), then switched by the coordinator in a user-approved cold-copy cutover (20:43:48Z stop, 20:44:37Z start, V65/V66 applied, fingerprint PASS, real Claude Code capture). The Codex /hooks trust step for the 7 changed commands is pending and the user's. - evidence/artifacts/sota-refresh-20260925/ai-memory/: the driver and the verifier's scripts byte-identical, sanitized step results, cutover.json with the coordinator's aggregate facts only (no cold copy, fingerprint file or store content), a read-only read-back and results.json with both arms, the verifier verdict and what was not published. - evidence/receipts/ai-memory-240-qualification-20260925.json (native_cli_e2e). Review repairs (2026-09-25, independent review needs_changes): the receipt adds the failed production compare check (claude_settings_unchanged: ~/.claude/settings.json changed at 19:01:59Z by an unidentified writer); results.json corrects the verifier's two wording defects in place and keeps the originals in a corrections list; recorded_at_utc is the time the receipt content was written. Final round: the #792 limitation says the fix is not in 2.4.0 and was released in v2.4.1 at 2026-09-25T21:46:32Z. manifests/evidence.json registration follows in the branch's last commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * mcporter 0.14.1 qualification evidence (sanitized artifacts, cutover facts and receipt) Qualified on nativestack-5975wx-20260925 against the installed 0.13.13: the verified npm tarball (sha256 8e489864..., SLSA provenance from release.yml at refs/tags/v0.14.1) in a new prefix, and the same 37-step acceptance in isolated daemon namespaces passing 37/37 on both versions with identical outcomes. An independent verifier agreed and measured the one difference the acceptance cannot see: without ps on the daemon's PATH, 0.14.1 cannot connect keep-alive servers and `daemon stop` refuses. The coordinator then relinked bin/mcporter with no daemon running; `list socraticode --brief --no-oauth` listed 26 tools and codebase_health reached Qdrant. - evidence/artifacts/sota-refresh-20260925/mcporter/: harness and verifier scripts byte-identical, both arms' results, the per-step raw outputs bundled into two JSON files (deduplicated), integrity and provenance outputs, cutover.json, a read-only read-back and results.json. - evidence/receipts/mcporter-0141-qualification-20260925.json (native_cli_e2e). Review repairs (2026-09-25, independent review needs_changes): the staging install's package.json and package-lock.json are published as stage/package.evidence.json and stage/package-lock.evidence.json (bytes unchanged), so they are not repository dependency manifests and the OSV inventory and Dependabot scope stay as they were; results.json records the renames; recorded_at_utc is the time the receipt content was written. Delta review (2026-09-26): the two npm registry key .pem files are listed under not_published (source https://registry.npmjs.org/-/npm/v1/keys), because .gitignore excludes *.pem; they were never committed. manifests/evidence.json registration follows in the branch's last commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * ai-memory 2.4.1 qualification and cutover evidence (sanitized artifacts and receipt) Qualified on nativestack-5975wx-20260925 against production 2.4.0 on restored copies of one online backup (2.4.0 arm 26/26, 2.4.1 arm 28/28: V66 -> V67 with nothing missing, 2.4.0 refuses the V67 copy, path-only install-hooks render, and the #792 A/B: 65/65 keepalive timers and all 60 vanished peers reclaimed on 2.4.1, none on 2.4.0), reviewed by an independent verifier (agree, 8 minor defects). The coordinator then performed the user-approved cutover on 2026-09-26: stop 00:14:12Z, cold V66 copy, relink, path-only hook rewrite, start 00:19:51Z with V67 applied, fingerprint PASS, keepalive on 2 of 2 live sockets and real Claude Code capture. The Codex /hooks re-trust is pending and the user's. The 2.4.0 receipt stays as history. - evidence/artifacts/sota-refresh-20260925/ai-memory-241/: the driver and the verifier's scripts byte-identical, sanitized results for both, cutover.json with the coordinator's aggregate facts only (no cold copy, fingerprint file or store content), a read-only read-back and results.json (steps for both arms, the verdict, and the verifier's wording corrections). - evidence/receipts/ai-memory-241-qualification-20260925.json (native_cli_e2e); recorded_at_utc is the time the content was written. Delta review (2026-09-26): results/step4-rehearsal.json has 28 checks, all true, so the receipt and the B-arm excerpt say 28 (the original 29 is kept in results.json corrections), and the R1/A17 control-build limitation cites issue #274. manifests/evidence.json registration follows in the branch's last commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * headroom 0.39.0 qualification evidence, not switched (sanitized artifacts and receipt) Qualified on nativestack-5975wx-20260925 against the pinned 0.37.0: the verified wheel in a new prefix (RECORD 555/555), the unchanged upstream #3736 test 6/6 (0.37.0: 4/6), and the stack's MCP path over stdio with byte-exact retrieval and cross-version store reads; an independent verifier agreed. The coordinator retains 0.37.0: at v0.39.0 the '[N lines omitted: ...]' notice counts levels over all input lines (log_compressor.py:461-487 and the Rust format_output; checked with gh), so fixture A's '[593 lines omitted: 1 ERROR]' reports the kept CRITICAL line as omitted; the Rust detector still lacks the #3736 timestamp guard; #299 pins 0.37.0 in the bootstrap and #296's counters run at it. No pin changes. - evidence/artifacts/sota-refresh-20260925/headroom-039/: harness and verifier scripts, key results as files, all other outputs and the production snapshots bundled, a read-only read-back and results.json (decision reasons, overturn condition, beacon note, corrections). - evidence/receipts/headroom-039-qualification-20260925.json (native_cli_e2e). manifests/evidence.json registration follows in the branch's last commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Linux pins, hook template, recipes and docs: ai-memory 2.4.1 and mcporter 0.14.1 adoption/pins-linux-x86_64.json moves ai-memory to the v2.4.1 archive (sha256 15cafdc4..., 16,026,020 bytes, equal to the release sidecar, the API digest and the release body; binary dec065a8...; tag commit 433a19f3) and mcporter to the 0.14.1 npm tarball (sha256 8e489864..., integrity sha512-/6pv...). The notes carry the forward-only migration and the ps-on-PATH requirement. pins-macos-arm64.json is unchanged: the Mac qualifies on its own. adoption/templates/claude.settings.template.json: the eight ai-memory hook commands name tools/ai-memory-2.4.1, matching the live source host after the upstream install-hooks rewrite. bootstrap.md, platforms/linux-wsl2.md and platforms/macos-arm64.md say the Linux pin file and the template changed after v2026.09.25.2, and tell a Mac to point the rendered hooks at its installed 2.3.2 prefix. recipes/README.md: both rows name the new pins, and a new "Upgrading an existing store" section gives the cold-copy procedure the host used. Copies follow in docs/token-efficiency-stack.json, tools/token-report/README.md and observability/native-data/README.md. The two catalogs/landscape/upstream-snapshot.json rows come from a fresh run of their gh api endpoints, and the saturation-audit rows cite the new receipts. catalogs/landscape/foundation.json winner pins do not move. tests/test_adoption_bootstrap_macos.py required every shared component to carry the same version and npm hash on both platforms, which a Linux-only move breaks. It now names each Mac-behind-Linux lag exactly (ai-memory 2.3.2/2.4.0 and mcporter 0.13.13/0.14.1, with their receipts), so a move on either side fails until the table is reviewed, and compares npm hashes only between equal versions. Review repairs (2026-09-25, independent review needs_changes): - recipes/README.md: the hook commands' 49374 is an example; use the host's own server URL (NativeStack binds 127.0.0.1:49474, and 49374 is another distro's default there) and record what NativeStack used. The upgrade section states the order hazard: bootstrap-linux.sh repoints bin/ai-memory at once, so an existing store is stopped and cold-copied before any bootstrap re-run; adoption/update.md step 1 says the same. - platforms/macos-arm64.md: the changed-after marker in step 3 states only the difference and points to a standing "ai-memory hook paths on macOS" note, which gives the full recipe command (--apply, --capture-mode allowlist, --no-capture-prompts, --server-url): with no stored mode, v2.3.2's installer falls back to denylist. - catalogs/landscape/foundation.json durable-memory: one appended, dated limitation for the 20:44Z cutover; the earlier 2026-09-25 text is kept, and the winner pin (a verdict field) does not change. Final round (2026-09-26): the ai-memory pin, the hook template, the recipes and docs, the upstream-snapshot row, the saturation-audit row and the Mac-lag test pair move from 2.4.0 to 2.4.1 after the user-approved 2.4.1 cutover (the 2.4.0 receipt stays as history). The upgrade section covers both cutovers and warns against `pgrep -f` in the quiesce wait. platforms/macos-arm64.md now says accurately that #299 also added ccusage, headroom, repomix, serena, socraticode and toon to the Linux pins, and which of the ten have Mac entries. The foundation.json limitation names both cutovers. Delta review (2026-09-26): the ai-memory install_note dates the three-value check to 2026-09-25 (22:43Z qualification, 23:10Z verifier) and says 2026-09-26 re-checked the API digest. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Decision records: correct the ai-memory 2.4.0 review; record the workstation SOTA refresh docs/decisions/2026-09-25-ai-memory-2-4-0-release-review.md gains a dated correction section; the original text is kept. GHSA-9pj6-vhgr-3mwh is not reachable on this host (stateless transport: GET /mcp 405 with allow: POST, no --http-stateful), and 2.4.0 writes no automatic pre-migration archive for a store created by 2.3.2 (snapshot_before_db_migration is gated to the 1.x -> 2.0 upgrade). It also records that the upgrade was performed on 2026-09-25 and cites the new receipt. docs/decisions/2026-09-25-workstation-sota-refresh.md is the dated research-convergence record for the whole workstation-lane refresh: per unit, the decision, its primary sources (releases, compares, advisories and the refuter's key finding), the alternatives and the overturn condition. rtk, markitdown, mcporter and ai-memory were qualified and switched; socraticode 1.15.0 is qualified with the cutover deferred to no earlier than 2026-10-01; dagu 2.17.2 is held by the trading lane; headroom stays 0.37.0 because 0.38.0 carries #3736 on the MCP compress path; qmd, qdrant, ccusage, worktrunk and vllm are already current; llama.cpp b11146 is recorded, not qualified; both models are retained; the ai-memory Nemotron embedder is evaluation-only. The research and qualification results are cited as private inputs only. Review repairs (2026-09-25, independent review needs_changes), each fact re-checked with gh api at about 22:35Z: - headroom: v0.39.0 (2026-09-25T19:29:55Z) contains d971f7c3 (compare ahead 86, behind 0; the release notes list #3748), so the overturn condition is met; qualifying 0.39.0 is the queued follow-up. - ai-memory: only commands/hook*.rs and commands/install_hooks.rs are unchanged in the v2.3.2...v2.4.0 compare (config.rs, serve.rs, run.rs, reindex.rs, mcp_bridge.rs and six ai-memory-core files changed). - ai-memory v2.4.1 (2026-09-25T21:46:32Z) carries the #792 fix and the forward-only V67 migration but not #859; it meets the "a 2.4.1 appears" condition, and qualifying it (another rehearsal and the user's approval) is queued. Recorded in both documents. - The embedder section cites issue #274 (the Mac-only harness, A17, "the workstation now owns R1") and catalogs/foundation/memory-stack-20260925.json; #274 covers the C3/C4 rerun, and the Nemotron-prefix evaluation comes after it. - The release review gains a dated one-line pointer under its title. - The rtk and markitdown receipts are linked now that #291 has merged. Final round (2026-09-26): - ai-memory: the record and the release review state that 2.4.1 was rehearsed and, with the user's approval, switched on 2026-09-26 (00:14:12Z stop, 00:19:51Z start, V66 to V67), replacing the "needs another rehearsal and the user's approval" sentences and "2.4.0 serves". - headroom: 0.39.0 was qualified and is not switched. The reasons are source-checked with gh at 00:28Z: the omission notice counts levels over all lines (log_compressor.py:461-487 and the Rust format_output; same Python blob since v0.37.0; #3635 changed only the Rust file), the Rust detector lacks the #3736 guard, and #299/#296 run 0.37.0. - The embedder section cites experiment.json and convergence.json for the C1 to C3 arms (the catalog has no C1 label). Delta review (2026-09-26): the refresh record says 2.4.1 arm 28/28 (results/step4-rehearsal.json has 28 checks), and the release review's "now name 2.4.0" is dated to the 2.4.0 cutover and marked superseded by 2.4.1. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Stack pins ai-memory 2.4.1 and mcporter 0.14.1; register evidence and regenerate reports Rebuilt on main c09dd6d (#306) per the hot-file protocol: the reviewed stack.json diff, the same registered path set and all four qualification receipts as the reviewed head b64f9b0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 26, 2026
… vLLM 0.30.0 embedder switch Mirrors the workstation's #296 method on wsl-authoring-20260923 (13 of #296's 16 tasks; 29 children, 15 Sonnet 5 + 14 Opus 5.5 at effort max). Publishes each tool's check result and independent verifier verdict, including the refuted rtk claim (branch -a '+' marker; the bare git show exclusion missing the -C form), Serena's missed cross-file references and repomix --compress dropping a def. Adds a baseline_kind field so only counterfactual_raw pairs read as reductions. Adds the laptop's vLLM 0.25.0 -> 0.30.0 qualification and switch receipt, reporting set identity, strict-order failures and the production-vs-production control totals. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 26, 2026
… vLLM 0.30.0 embedder switch (#316) Mirrors the workstation's #296 method on wsl-authoring-20260923 (13 of #296's 16 tasks; 29 children, 15 Sonnet 5 + 14 Opus 5.5 at effort max). Publishes each tool's check result and independent verifier verdict, including the refuted rtk claim (branch -a '+' marker; the bare git show exclusion missing the -C form), Serena's missed cross-file references and repomix --compress dropping a def. Adds a baseline_kind field so only counterfactual_raw pairs read as reductions. Adds the laptop's vLLM 0.25.0 -> 0.30.0 qualification and switch receipt, reporting set identity, strict-order failures and the production-vs-production control totals. Co-authored-by: seathatflowsinourveins <234074349+seathatflowsinourveins@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 26, 2026
…ecks, four real limits, and a correction to #296's QMD row (#343) * Record the token stack inside native Codex sessions, and correct #296's QMD row Codex side of #296: one native `codex exec` session per tool (codex-cli 0.155.1, gpt-6-astra, effort max) in a dedicated worktree, using RTK's standing instructions, the Context Mode plugin, the #289 user-scope MCP servers and the CLIs. 11 of 15 tools pass their answer checks; RTK +17782 saved for the run-only worktree; nine exact o200k comparisons from 28.2% to 99.8%; Codex-returned usage 7,085,015 input (6,253,312 cached) and 164,906 output over 17 sessions. The four failures are real limits: jCodeMunch is project-scoped (#240), QMD's catalog index excludes adoption/, Repomix --compress dropped a declaration, Context Hub lacks the upstream facts. #296's QMD PASS retrieved a wrong document for its question; its README now says so. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Sanitize the Codex E2E receipt after cross-family review (PR #343) Codex review (gpt-6-astra, read-only) confirmed every number and evidence class and asked for three fixes: replace the live-clone path, the session-derived codebase-memory project id and the SocratiCode collection id with placeholders, and state the QMD document count per attempt (121, then 122 as the live index grew). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 27, 2026
…asks, metrics, gate, procedure) Freezes the E2E preregistration half of full-save/PLAN.md step 9 (the baseline-receipt half already merged via #369) before any organic run: 77 frozen task definitions (16 reused Claude + 15 reused Codex information needs from #296/#343's receipts, 46 seeded/control tasks), every AA §8.3 metric/threshold/guardrail/outcome rule with its three "proposed" bounds now stated as frozen, the AA §8.1b capability gate, the AA §8.4 procedure, the AA §8.5 overturn conditions, the identity scheme, the Q1 blind-process qualification, the Workflow script that will run arms B/A/A0 (syntax-checked only, never executed), and the measurement runbook. Corrected on independent Claude review (11 findings, 95 turns) and a GPT-6 cross-family review (2 findings, 1 high) plus one combined repair round (13/13 addressed with failing-first evidence): relocated from blueprints/token-practice-e2e/ to evidence/artifacts/token-adoption-e2e-20260926/ per adoption-audit/PLAN.md line 402; moved the Workflow script out of examples/claude-native/workflows/ to avoid that directory's shared PACKET/ROUTING contract test, which this one-off frozen script correctly does not match; corrected PR-A's status (its base branch merged as #369 -- baseline only, tooling-fix scope still open); added the six AA §6 machine-readable policy fields the JSON was missing; fixed a wrong Codex arm-A lifecycle binding; blocked the one information need (an otel-tui receiver task) that no current role may legally run, pending a named permitted role; froze the coordinator's own main-dispatch task at xhigh instead of max; fixed a denylist word-boundary gap that let underscore-qualified MCP names slip past; and sealed the five normative files' SHA256 with an append-only amendment rule. Also refreshes PR #376 and #364's status to merged (both landed while this unit was in review), while recording that #364's field-preservation gap (workflow.run_id/tool_use_id still missing from collector.yaml) persists after its merge. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 27, 2026
…asks, (#381) metrics, gate, procedure) Freezes the E2E preregistration half of full-save/PLAN.md step 9 (the baseline-receipt half already merged via #369) before any organic run: 77 frozen task definitions (16 reused Claude + 15 reused Codex information needs from #296/#343's receipts, 46 seeded/control tasks), every AA §8.3 metric/threshold/guardrail/outcome rule with its three "proposed" bounds now stated as frozen, the AA §8.1b capability gate, the AA §8.4 procedure, the AA §8.5 overturn conditions, the identity scheme, the Q1 blind-process qualification, the Workflow script that will run arms B/A/A0 (syntax-checked only, never executed), and the measurement runbook. Corrected on independent Claude review (11 findings, 95 turns) and a GPT-6 cross-family review (2 findings, 1 high) plus one combined repair round (13/13 addressed with failing-first evidence): relocated from blueprints/token-practice-e2e/ to evidence/artifacts/token-adoption-e2e-20260926/ per adoption-audit/PLAN.md line 402; moved the Workflow script out of examples/claude-native/workflows/ to avoid that directory's shared PACKET/ROUTING contract test, which this one-off frozen script correctly does not match; corrected PR-A's status (its base branch merged as #369 -- baseline only, tooling-fix scope still open); added the six AA §6 machine-readable policy fields the JSON was missing; fixed a wrong Codex arm-A lifecycle binding; blocked the one information need (an otel-tui receiver task) that no current role may legally run, pending a named permitted role; froze the coordinator's own main-dispatch task at xhigh instead of max; fixed a denylist word-boundary gap that let underscore-qualified MCP names slip past; and sealed the five normative files' SHA256 with an append-only amendment rule. Also refreshes PR #376 and #364's status to merged (both landed while this unit was in review), while recording that #364's field-preservation gap (workflow.run_id/tool_use_id still missing from collector.yaml) persists after its merge. Co-authored-by: Scout <scout@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 27, 2026
…ta from the 2026-09-27 verdict wave (#401) * E1/E2 token E2E receipts: dated adjudications, errata and reproduced answer oracles Applies the 2026-09-27 cross-family GPT-6 review of the E1 (#296) and E2 (#316) receipts, append-only: every recorded field keeps its value, and each receipt keeps its own serialization. - Every tools[] row gains adjudication {as_of, task_acceptance, bad_input_control_recorded, basis, evidence}, using the review's status table and vocabulary (partial, retracted_later, untested, unsupported). - E1: QMD retracted_later (wrong_document, corrections_to_296); Repomix retracted_later (count=47 PASS on a 48-function file; repomix v1.18.1 PythonParseStrategy.ts L81-86 tests only a def's first line); ai-memory untested (base_empty=True is a vacuous pass); the other thirteen partial with no recorded failing control. A dated errata block names the pre-errata sha256 the preregistration records as a historical identity. - E2: context-mode, jcodemunch, qmd and ast-grep stay partial with no recorded failing control; SocratiCode's comparison counts an abridged transcription; RTK retracted by its verifier; Repomix's FAIL accurate. - READMEs carry dated corrections; the review is retained, sanitized, beside E1. - tests/test_token_e2e_receipt_checks.py: receipt contracts (red before the errata) and answer oracles rebuilt from pinned Git blobs, hash-matched to the recorded baselines, each with failing controls (red under a weaker historical-style check). Sources: docs/acceptance-evidence-policy.md (Discriminating controls); https://github.com/yamadashy/repomix/blob/v1.18.1/src/core/treeSitter/parseStrategies/PythonParseStrategy.ts#L81 evidence/artifacts/token-e2e-codex-20260926/receipt.json#/corrections_to_296. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Token layer: routing rules and recipe limits for known upstream defects Closes the documentation, recipe and routing gaps from the 2026-09-27 per-tool verdict wave for the token-efficiency layer. Every rule cites the pinned upstream source or a committed receipt; no pin changes. recipes/README.md (token-tool rows and sections): - agentsview: v0.44.0 unqualified; worker populations need --include-children, --include-automated and --include-one-shot, and --fts reads message bodies only (v0.43.0 docs/session-api.md, docs/commands.md); archive answers are observation. - ast-grep: every 0.45.3 command registers customLanguages from any sgconfig.yml in the working directory or a parent (lib.rs L107-139, config.rs L98-131/L280-299); PR #2960's opt-in is unreleased; outline (#2957) and YAML rule (#2963) limits. - ccusage: token-only while any model is unpriced (v20.0.24 json-output.md Unpriced Models; config-files.md pricingOverrides). - codebase-memory: trace_path excludes tests and evidence by default, search_code hits can be mentions (v0.11.0 src/mcp/mcp.c schemas); caller-set procedure. - context-hub: check metadata.versions (FastAPI guide 0.136.3 vs 0.141.1). - context-mode: npm tarball digest and plugin commits are separate identities. - headroom: v0.39.1 changes only the proxy limiter, so the omission-count defect stands. - markitdown: HTML as .txt passes through; -x html / -m text/html; selective extras. - mcporter: --name does not select the server; a positional token becomes the selector (v0.14.1 call-arguments.ts L164-192, ephemeral-flags.ts L97-105); the jCodeMunch example gains --server. - qmd: MCP query expands and reranks by default (v2.8.3 README, server.ts); typed lex + rerank:false + bounded get; catalog scope; refresh guard for #989/#991 (src/store.ts reindexCollection). - repomix: multi-line Python signatures are dropped (v1.18.1 PythonParseStrategy.ts L81-86); exact-definition tasks go to an uncompressed pack or the original. - toon: BOM root-string round trip (#339); no record-count threshold upstream; nested-uniform columns are tabular (tabular.ts L61-78). docs/token-session-handbook.md: a "Known upstream limits behind the lanes" section outside the carrier text (every carrier line unchanged), and row pointers. docs/token-practice.md: counts, comparisons and acceptance rules (invocation counts are not success rates; comparisons need task acceptance; count recovery reads; gpt-tokenizer's special-token contract; receipts carry adjudications). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Apply the GPT-6 review of this branch: exact inventory oracle, TOON eligibility wording One cross-family review round (GPT-6 through codex-omniroute, read-only) returned DEFECTS FOUND with two findings; both are fixed: - P2 tests/test_token_e2e_receipt_checks.py: inventory_oracle checked only the claimed count plus subset membership, so a correct count with missing, duplicated or no names passed. It now requires the exact name set with one entry per function. New controls keep the correct count while dropping, duplicating or inventing a name, or listing none; they failed against the old oracle first (red) and pass now. - P3 docs/token-session-handbook.md (and the matching recipes/README.md toon row): "accepts any non-empty uniform array" overstated TOON 4.1.1's tabular eligibility. The encoder needs non-empty objects sharing one key set, with columns of primitives or, recursively, of non-empty objects sharing one key set; an empty object or an array-valued column falls back to list form (packages/toon/src/encode/tabular.ts L6-78 at v4.1.1; confirmed with the installed 4.1.1 CLI on [{}], [{"a":[]}], [{"a":{}}] and [{"a":{"b":1}}]). Also: the ast-grep oracle keys counts by path under scripts/ rather than file stem, and E1's Context Mode adjudication states that attempt 1's marker-grep FAIL on a refused call is not shown to be a run of attempt 2's answer check (the receipt does not record that command). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Apply the independent verifier's findings: TOON source line, citation anchors, E2 Serena erratum The independent verifier of this branch reported five defects; all are fixed: - docs/token-session-handbook.md, the TOON Sources line under the carrier (on main before this branch) still said "The encoder accepts a non-empty uniform array", the wording GPT-6 finding P3 refuted. It now states the encoder's rule from packages/toon/src/encode/tabular.ts L6-78 at v4.1.1 and links the handbook's "Known upstream limits behind the lanes" section. It is not a "- " carrier line, so tests/test_token_lanes_subagent_start.py's verbatim-line check is unaffected. - Citation anchors: SocratiCode's INCLUDE_DOT_FILES default is documented under README.md "Indexing Behaviour" (v1.14.0 L1574-1579), not "Ignore Rules", so the handbook links #indexing-behaviour. codebase-memory-mcp v0.11.0 search_code gets its own link to src/mcp/mcp.c#L632-L659 in the handbook, and the recipe gives index_repository (L466-481), trace_path (L532-560) and search_code (L632-659) one link each. Both upstream files matched the retained copies by sha256 on 2026-09-27. - E2 receipt errata.items[4] and the E2 README erratum now name Serena's accurate reference FAIL (find_referencing_symbols missed 6 of 8 call sites; the definition check passed) and give each group of partial rows its own reason. Only that one finding string changed in receipt.json (same serialization, 1 line); written_at is unchanged because this refines the unmerged 2026-09-27 erratum on the same day. - tests/test_token_e2e_receipt_checks.py: the require_commit docstring now says the guard is modelled on RetainedEvidenceTests' Git check and adds the pinned-commit probe (tests/test_adoption_status.py L1721-1724 checks only git and the checkout). Sources: toon-format/toon v4.1.1 packages/toon/src/encode/tabular.ts; giancarloerra/SocratiCode v1.14.0 README.md; DeusData/codebase-memory-mcp v0.11.0 src/mcp/mcp.c; the E2 receipt's own serena row (tools[5]) and README. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Scope
c303d8e3(currentorigin/main, after Codex MCP servers at user scope so every worktree inherits the token tools; --pinned-versions coverage check #289)lane:foundationevidence/artifacts/token-e2e-ultracode-20260925/(README and receipt.json)evidence/hosts/nativestack-5975wx-20260925/(ast-grep, codebase-memory-mcp, context-hub, agentsview, otel-tui)docs/token-efficiency-stack.mdmanifests/evidence.jsonregistrationsWhy. Installing and configuring a tool does not show a saving. Each claim here pairs a subagent actually using the tool, with its answer checked against the plain baseline, with that tool's own counter where it has one and an exact
o200k_basecomparison of the same information need.Host changes made before the run
These are not part of this diff.
manifests/stack.jsonpins, checked against upstream sha256 or npm integrity: ast-grep 0.45.3, codebase-memory-mcp 0.11.0, Context Hub 0.1.4, agentsview 0.43.0, otel-tui 0.7.5.headroom(headroom mcp serve --proxy-url http://127.0.0.1:1, so no model traffic is rerouted),codebase-memory(with its web UI turned off:ui_enabled=false) andqmd(qmd --index native-agent-stack-catalog mcp). Codex MCP servers at user scope so every worktree inherits the token tools; --pinned-versions coverage check #289 mirrors these in Codex.gpt-tokenizer@3.4.0.Evidence-class table
native_provenon this hostreceipt.jsontools[]local_integration(exacto200k_basecounts of retained artifacts; not provider savings)tools[].exact_comparisoncounters.rtk_gain_e2e_worktree_onlycounters.headroom_savings_lifetimeworkflow_consumption(child-usage.mjs)claude -psession loads the three new servers and calls one tool on eachnative_provenfresh_native_sessionnative_proven(self review only)evidence/hosts/…Retained failures and gaps
The full list is in the README.
codex execreturns the usage limit until 2026-09-30 23:50.set(...)asenvironment_dump(a false positive)..htmlextension; on.txtit passes the input through.Local commands run
uv run …stands foruv run --no-project --python 3.13 --with-requirements .github/requirements-ci.txt python.Decision record
None. This PR is evidence only and makes no pin or verdict change.
Host evidence
host_receipts.py validatepasses.independent_sessionverdicts. Codex is unavailable until 2026-09-30.platform_statuschange is made by hand.Checklist
🤖 Generated with Claude Code