Repository navigation
Token evidence: RTK coverage study and 18 token-tool currency receipts (sanitized, revision-bound) - #409
Merged
Conversation
Preserve the historical RTK aggregate tables, synthetic exactness output and method, with private sample fragments removed and a dated M-R1 interpretation erratum. Publish 18 source-review currency records with live release metadata, explicit evidence classes and retrieval dates. Preserve scratch/current differences, including Headroom 0.39.1 and the context-hub PR #182 correction. Sources: rtk-ai/rtk v0.50.0 (1d87b8e719ce0a50c223cd93ca64dd16921f9aec), src/main.rs, src/hooks/decision.rs, src/discover/mod.rs and registry.rs; rtk-ai/rtk dev-0.51.0-rc.467 (a89a31494670fcec8ffa20d939dd94c64bd998fb). The 18 upstream repositories, pins and finding URLs are retained in evidence/artifacts/token-tool-currency-20260927/records/*.json. Release retrieval: https://docs.github.com/en/rest/releases/releases#get-the-latest-release and https://cli.github.com/manual/gh_api. Publication tests: https://docs.python.org/3/library/unittest.html. Leak scan: gitleaks/gitleaks v8.30.1 README.md. Validation: 5 offline unittest checks passed; native Gitleaks and identifier scans found no leaks. scripts/validate.py reports only the 25 new evidence files awaiting the coordinator's hash registration. No pins or runtime configuration changed. Recommended label: lane:foundation Content written by GPT-6 (gpt-6-astra, effort max) through the Codex worker lane and the OmniRoute pool; committed by the coordinator harness. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Attribute RTK T7 to rc.467, qualify query timestamps and Serena's main distance, correct all metadata coordinates against repository commit 1c32ad2, and add discriminating unittest controls. Preserve historical output bytes and scratch records. Sources: rtk-ai/rtk v0.50.0 (1d87b8e719ce0a50c223cd93ca64dd16921f9aec), dev-0.51.0-rc.467 (a89a31494670fcec8ffa20d939dd94c64bd998fb), Cargo.toml:3; full-save/rtk/REPORT.md:37-41,193,205,270 and exactness.sh; Python unittest https://docs.python.org/3/library/unittest.html; Gitleaks v8.30.1 README; docs/acceptance-evidence-policy.md and docs/lanes.md. Refresh the PR evidence to distinguish reviewed a0ac1cde registration and passing harness logs from the unregistered b0fc8f5c repair checkout. Leave manifest registration, report generation and commits to the harness. Content written by GPT-6 (gpt-6-astra, effort max) through the Codex worker lane and the OmniRoute pool; committed by the coordinator harness. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ch stamps and revision scope Coordinator corrections after the independent repair verification: test_pin_metadata_lines_contain_the_component_version now reads each cited file with git show <metadata_revision>:<path> (the records cite coordinates at 1c32ad2, which later changes may move); METHOD.md describes the two retrieval batches separately and no longer claims the revision's files match every later tree. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
force-pushed
the
claude/w3-evidence-20260927
branch
from
September 27, 2026 22:23
1088879 to
3fda98b
Compare
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…inate (G1)
The nested lines_at helper turned every failed git show into skipTest, so a
nonexistent metadata path or revision bypassed provenance validation. The
module-level metadata_lines helper follows tests/test_release_pin_contents.py
(git cat-file -e <rev>^{commit}; fail in CI, skip locally): an absent revision
skips only when GITHUB_ACTIONS is not "true" (validate.yml checks out full
history, fetch-depth: 0), and a failed git show at an existing revision fails
with the revision, path and git's first stderr line.
Discriminating controls (red first against the old skip: FAILED (failures=3),
'skip' != 'fail'): a bogus path at HEAD fails with and without GITHUB_ACTIONS,
and a missing revision skips locally but fails with GITHUB_ACTIONS=true.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d report; record the bucket recount (G2, G3) G2: the transcript-file, subagent and window figures appear in no published aggregate; a lead-in now says the unpublished REPORT.md states every figure in that paragraph and that only the 59,041-call total is also retained in tables.txt and frame-summary.txt. The four original sentences are unchanged. G3: a scripted recount of the 22 bucket rows in frame-summary.txt gives 15,623 missed parts (matching TOTAL and defer+deny+NO_ROW) and 25,947 covered parts, 7 fewer than the TOTAL line and ask row (25,954). One dated sentence records both sums and that the published files do not attribute the difference. The data files are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ain observation (G4) Of the nine batch-a records, context-mode, ccusage, qmd, repomix and toon state current-main, commits-ahead or unchanged main distances with no unreleased_main block and no returned head of main. context-mode (10 -> 12) and ccusage (183 -> 198) retain counts from compares ending at fixed commits, with nothing showing those commits were main's head; qmd, repomix and toon name a compare of the mutable main ref whose returned head and count were not retained, so their only retained distances are the 2026-09-26 scratch_record figures. One dated limitation paragraph names them and requires re-querying before use. headroom and rtk make no main-distance claim; markitdown and serena already bound theirs. No record changed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eview Hot-file protocol (docs/lanes.md): last commit only. register_file refreshes the hashes of tests/test_token_full_save_evidence.py and the two METHOD.md files; component_matrix.py --write and new_host_grand_list.py --write produced byte-identical reports. validate.py, evidence_manifest.py --check and component_matrix.py --check pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…vations each record retains The README said the records retain the observed heads, but only jcodemunch-mcp, codebase-memory-mcp and agentsview keep an unreleased_main head; the other five keep a 2026-09-26 scratch head and compare end commits. The METHOD limitation now names those per record. No number changed. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ed for the wording fix (manifests/evidence.json only) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…t, branch files re-registered) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…t, branch files re-registered) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Owner
Author
|
Review record, 2026-09-29 (native-agent-stack-10). Merge candidate: head
|
4 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Publish the two scratch-only PR-EV results so their aggregate numbers, source
references and release observations can be inspected from the repository:
the historical RTK coverage study and all 18 token-tool currency records.
Changes
sanitized exactness.out, method and a records table. Preserve numeric
outcomes except for the removed unsupported-key list, including its counts;
remove the local recall handle. This repair leaves all three data files unchanged.
records, a method and a records table. Retain reviewed pins, scratch claims,
actual allowlisted live release metadata, channel-specific behind_by values,
primary-source findings and explicit retrieval dates.
classification is not the later M-R1 eligible-part gate.
unreleased-source distances and the context-hub Secret-storage practice: per-host 0600 store, agent read guard hook + deny rules, managed-profile install, names-only checker #182 closed/unmerged correction.
fields, preserved RTK outcomes, evidence boundaries and identifier patterns.
cargo-test boundary and remove the unretained installed-version/help claim.
1c32ad2; add component/version and revision
controls for all 18 records, including secondary pins.
distance not re-verified, and label publication checks structural validation.
synthetic samples; retain failure and passing observations separately.
Evidence (with classes)
611-observation aggregate classification; no new private population replay.
output; not a new RTK run and not unchanged upstream tests.
for all 18 tools on 2026-09-27, selected package-channel checks, tagged/commit
source review and finding-specific primary URLs.
failed before receipt corrections (10 tests, 56 subtest failures, exit 1).
The fresh passing run returned 10 tests OK, exit 0. An intermediate run
failed twice because the new control used candidate names instead of
landscape winners'
component_id; reading the original entries corrected it.These are artifact consistency checks, not native tool execution or adoption.
The build's older missing-artifact failure had no retained stdout and did not
demonstrate identifier detection.
publication artifacts rejected each planted email, UUID, personal path,
bearer, credential-shaped value and session identifier. The unwrapped check
returned exit 1 with six failures; the permanent unittest requires each
rejection. No sample values are published in these receipts or logs.
committed publication bytes; tables.txt still has the original aggregate
SHA256. All 18 scratch records, pin versions and latest-release snapshots are
unchanged. Eight metadata files were read from the named repository revision
and byte-compared with the checkout.
rtk git diff --checkreturned 0.directories and the unittest module are retained below. The shared pattern
scan checks all 25 public artifacts. Neither scan establishes detection of
every sensitive string. No credential stores were read.
installations/upgrades, upstream test reruns or fresh tokenizer parity trials.
Retained repair outputs, 2026-09-27: the covering command is
rtk python3 -m unittest tests.test_token_full_save_evidence, preceded byexport TMPDIR=/var/tmp/claude-evidence, at the worktree root. Start2026-09-27T14:12:05.239823+00:00, end 14:12:05.437827+00:00; exit 0,
stdout empty, stderr:
The red publication run returned
Ran 10 tests in 0.039sandFAILED (failures=56). The separate planted-control call usedrtk python3 -with the same TMPDIR and shared assertion; exit 1, stdout empty, stderr
excerpts (synthetic fixture, expected rejection):
Acceptance of the final branch (local integration, this host, 2026-09-27):
Ran 854 testsandOK (skipped=73).validate.shreported FAILS=0. The full suite ran 6603 tests with failures=1:tests.test_secret_path_guard.test_host_profile_copy_is_verbatim, which fails on branches older than main d022295 because this host installed Claude harness settings: credential-store and destructive-git denies, Opus role agents without frontmatter isolation, rtk-aware guard #402's guard at 14:04Z.test_pin_metadata_lines_contain_the_component_versionread each cited file at the record'smetadata_revision(git show <revision>:<path>), so later line moves on main cannot break it. It also reworded the two METHOD.md passages the verification found inaccurate.validate.shFAILS=0;tests.test_token_full_save_evidenceran 10 tests,OK.Fresh Gitleaks returned output: installed
rtk gitleaks versionreturned8.30.1. For each target below, the command at the worktree root wasrtk gitleaks dir TARGET --config .gitleaks.toml --redact --no-banner --no-color.Every command returned exit 0 and empty stdout; returned stderr follows.
Only terminal color escapes were stripped; times are the client's local times.
SOTA sources
The source selection reuses maintained native release APIs and upstream
implementations; no new runtime mechanism is introduced.
1d87b8e719ce0a50c223cd93ca64dd16921f9aec: src/main.rs,
src/hooks/decision.rs, src/discover/mod.rs, src/discover/registry.rs.
commit a89a31494670fcec8ffa20d939dd94c64bd998fb: historical fixture comparison;
Cargo.toml:3
confirms the 0.49.0 binary label. The supplied RTK REPORT.md:37-41,193,205,270
establishes the tested revisions, jq column attribution and cargo-test omission.
and gh api: supported metadata retrieval.
existing tests/test_token_e2e_preregistration.py artifact-contract pattern.
supported directory/file scan.
evidence/artifacts/rtk-exclude-widen-20260926/README.md: evidence boundaries and
records-table style. The full-save plan sections 2, 3.1 and 4.1 define scope.
sources named in every record, including manifests/stack.json, platform pin
files, the landscape winners and tokenizer sources. No pinned version changed.
Review dispositions
All nine findings are accepted and repaired; none is disputed.
The dated erratum cites pinned Cargo.toml and preserves exactness.out bytes.
and are opened by a unittest that checks component and version, including
secondary pins. Scratch records retain their original values.
numeric claim now excludes the removed unsupported-key list and its counts.
validation, which proves artifact consistency only.
the publication scan; their failing observation is retained beside the green run.
retrieved_atis the latest-release query's batch start; package/compareobservations have only a retrieval date. The prior ENOBUFS rerun is disclosed.
scratch_record; fixed-SHA comparison URLs remain source evidence only.
statement points to this PR Evidence section instead of private unit notes.
Review round (2026-09-29, native-agent-stack-10)
Head reviewed:
e3678505(main merged atba31dcd0), by two model families, read-only, on frozen diffs.gpt-6-astra, effort max,codex exec -s read-only; the model is the-mrequest, not an observed resolution)Repairs, each checked by an independent Opus stack-verifier that re-ran the acceptance commands and proved red-first on the previous source:
tests/test_token_full_save_evidence.pyhad turned every failedgit showinto a skip, so a nonexistent metadata path or revision bypassed validation. A module-levelmetadata_linesnow fails when the revision exists but the path does not, and fails in CI (GITHUB_ACTIONS=true) when the revision is missing (validate.yml checks out full history); it skips only outside CI. Two control tests fail against the old logic. Result: 12 tests OK, 0 skips, with the variable unset and set totrue(the second run simulates only the variable, not a CI run).rtk-coverage-study-20260927/METHOD.mdnow says the unpublished original report states the 1,945 / 1,606 / 1,594 / 53,895 figures and the window bounds; only the 59,041-call total is also retained in the published aggregates.token-tool-currency-20260927/METHOD.mdnames the five records (context-mode, ccusage, qmd, repomix, toon) whose 2026-09-27 figures have no retained main observation, and the README no longer says the records retain the observed heads for them. Only jcodemunch-mcp, codebase-memory-mcp and agentsview keep anunreleased_mainhead. No number changed.Final head
39e39bcbis0f0a67faplus a merge of mainb1be50c8by the hot-file protocol (26 files re-registered). Measured 2026-09-29 on it:python3 -B -m unittest tests.test_token_full_save_evidence12 tests OK; the three registry tests OK;scripts/validate.py,scripts/evidence_manifest.py --checkandscripts/component_matrix.py --checkexit 0 (run by the merge script).Evidence class of the repair: local integration checks measured on 2026-09-29 (
unittest,validate.py,evidence_manifest.py --check,component_matrix.py --check, the three registry tests with zizmor on PATH). They are not upstream tests. The RTK data files (tables.txt,frame-summary.txt,exactness.out) and all 18records/*.jsonare byte-identical to the reviewed head.Residuals and not done
metadata_revisionis 1c32ad2. Their line coordinates are valid at that revisionand may differ on main after this lands; the unittest reads the recorded revision.
retrieved_atprovenance differs by batch (METHOD.md): batch a shares batch stamps; batch b's per-recordstamps are retained, but how they were produced is not.
intentionally unpublished; full historical population regeneration is not
possible from aggregate receipts alone.
Serena's and MarkItDown's historical main-distance claims are not re-verified.
The command producing the T7 appendix was not retained.
all retained publication snapshots remain dated 2026-09-27.
changes only by re-registration of this branch's files in the last commit,
which keeps lane:foundation under docs/lanes.md:145-148.
datekey; the test'sIDENTIFIER_PATTERNSare not a subset ofscripts/validate.pyPRIVATE_CONTENT(each has patterns the other lacks); the transcript counts have no retained output beyond the attribution wording;rtk-coverage-study-20260927/README.mdL6 repeats the measurement window without the attribution; a revision that exists but is not a commit object is treated as absent bymetadata_lines.Recommended lane label: lane:foundation.
🤖 Generated with Claude Code