Skip to content

Split readiness G6 and record G5 closure criteria - #890

Merged
seathatflowsinourveins merged 1 commit into
mainfrom
codex/ns2604-readiness-g6-split-20261008
Oct 9, 2026
Merged

seathatflowsinourveins merged 1 commit into
mainfrom
codex/ns2604-readiness-g6-split-20261008

Conversation

@seathatflowsinourveins

@seathatflowsinourveins seathatflowsinourveins commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Scope

The dated readiness manifest still reports G6 as HOLD after #881 landed and the F9 receipt accepted G6a. Split that gate into G6a (MET), G6b (ongoing settings ownership), and G6c (ongoing records, #885/#886), and bind G5's six closure requirements without granting G5 acceptance. Update the matching R&D and follow-up state displays from the same accepted receipts.

  • Base commit: aba02ec3456d383bcc2fc72883db098f9be7918a
  • Lane: lane:foundation
  • Owned paths touched: tools/north-star/sources.json, catalogs/north-star/readiness.json
  • The existing builder regenerates the complete dated view, including current main-input provenance digests. Its 20 layer scopes and 36 SDK items remain. The evolving Monitoring status source matches the owner selector twice, so that single field becomes UNVERIFIED; its bounded snapshot is distinct from the separately preserved Monitoring follow-up.

SOTA sources

  • Existing reference implementation: native-agent-stack at aba02ec3456d383bcc2fc72883db098f9be7918a, tools/north-star/build_readiness.py, Receipts.claim and the supported --write/--check interface. This change uses the shipped receipt adapter and its digest-bound selectors. The gate-refresh procedure is documented in tools/north-star/README.md, Refresh on a gate change.
  • The original adapter's retained upstream survey is research/fullspeed-20261008/ns-readiness-manifest/upstream-research.md under the durable state root; the implementation follows Python 3.13 re for exact text selection and SLSA v1.2 verification for the distinction between input identity and acceptance. This is a source-binding refresh of that accepted implementation.
  • Gate evidence: coordination/command-center/windows/f9-g6a-881-20261008/RECEIPT.md, SHA256 24906eaf9cc3f041ed6bd500dd41095e4282a551ad1403d80a76a709ce7cf6a1; line 3 identifies #881's merge commit and timestamp, line 5 records F9 application, lines 14–15 record both fresh-session checks, and line 17 adjudicates G6a and the remaining G6 scopes. Installed gh 2.102.0 independently returned the same Make proactive harness improvement explicit in the shared core #881 merge metadata and OPEN draft record PRs #885 and #886.
  • G5 requirements: coordination/command-center/ITEM-ns2604-coop-20261008T210228Z.md, SHA256 d3868b6a0a81826407ade426c6f8a2c6511401177d3801eb24b3ae0e9023aa43, lines 16–21. The six criteria cover population witnesses; row evidence with conservative PENDING dispositions; the mined-list census; validation of the full 368-star/45-field/2-document asset; both-family reads of 113 action rows plus seeded 59-row samples per stratum with zero defects; and a cued release of the hash-pinned asset. These fields record requirements rather than passed checks.

Evidence-class table

Claim Evidence class Command / receipt
G6a is MET, #881 merged at 2026-10-08T20:41:37Z as e48de2e6fa1dd3b740af4793ea5824ea35476dba, and F9/fresh-session checks are retained source_review F9 receipt at the SHA256 above; bounded gh pr view 881 --json number,title,state,isDraft,mergedAt,mergeCommit,url
G6b/G6c continue, with #885/#886 covering records source_review F9 receipt line 17; bounded GitHub PR metadata
G5 has six required closure conditions and remains the open START gate source_review G5 task record at the SHA256 above, lines 12 and 16–21
The manifest reproduces the declared sources and joins each new gate field to exactly one retained selector local_integration Builder --write, --check, and focused gate/receipt integrity assertions, rc 0
Adapter and page-publication regressions pass synthetic tests.test_north_star_readiness, 32 fixture tests, rc 0
Repository publication integrity and scope pass local_integration python3 scripts/validate.py, rc 0; 70 components, 10,900 hashed files, 4 profiles, 239 receipts

Local commands run

nice -n 10 ionice -c2 -n7 python3 -I tools/north-star/build_readiness.py --write
# rc 0; 8 gates, 20 layers, 36 SDK items
nice -n 10 ionice -c2 -n7 python3 -I tools/north-star/build_readiness.py --check
# rc 0; manifest SHA256 0e2dba792fff6d4adbc13b47995474591f1024e8253bd7e5caf3a5c4db619aae
timeout 600 nice -n 10 ionice -c2 -n7 python3 -m unittest tests.test_north_star_readiness
# rc 0; 32 tests, OK
nice -n 10 ionice -c2 -n7 python3 scripts/validate.py
# rc 0; integrity and scope passed
git diff --check
# rc 0

Durable logs under research/fullspeed-20261008/ns-readiness-manifest/:

  • g6-split-targeted-tests.log, SHA256 fcee3223b60efc73e55954ea5d703f2b75952d3a6b8d9965e7127379f70b03cc.
  • g6-split-publication-validate.log, SHA256 9106406d71891d803b9ed1569f54c5df6195c29846e57e0f1c284e993ce92652.

Decision record

This is the accepted gate-receipt refresh requested by task task-ns2604-coop-20261008T212726Z. The two named immutable source records above establish its scope; source selectors retain full digests and exact line locators. A replacement adjudication receipt would supersede these values through the same supported refresh procedure.

Host evidence

This PR changes the readiness view and its selectors. Host-evidence receipt requirements are not applicable to its owned paths.

Checklist

  • New/changed GitHub Actions are pinned to a full commit SHA with a version comment (no floating tags). Not applicable to this data refresh.
  • New/changed workflows declare top-level permissions: {} and grant each job only what it needs. Not applicable to this data refresh.
  • No secrets are printed, logged or committed; no new required secret was added without a documented owner.
  • No new paid hosting, subscription or billing surface was introduced.
  • Peer-owned untracked files and worktrees were preserved (not deleted, moved or overwritten).

@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Oct 8, 2026
@seathatflowsinourveins
seathatflowsinourveins marked this pull request as ready for review October 8, 2026 22:45
@seathatflowsinourveins
seathatflowsinourveins force-pushed the codex/ns2604-readiness-g6-split-20261008 branch from b573e87 to 65d2259 Compare October 9, 2026 00:02
@seathatflowsinourveins
seathatflowsinourveins merged commit 72e1f94 into main Oct 9, 2026
49 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the codex/ns2604-readiness-g6-split-20261008 branch October 9, 2026 01:13
seathatflowsinourveins added a commit that referenced this pull request Oct 9, 2026
…894)

### Scope

- What this PR changes, in one or two sentences: adds `claude-pr-review.yml`, an on-demand, read-only review of one same-repository pull request head by `anthropics/claude-code-action` v1.0.247. A maintainer dispatches it from `main` with the pull request number and the exact head commit; the review goes to the job summary and nothing is posted to the pull request. It runs only while the repository variable `CLAUDE_PR_REVIEW_ENABLED` is `true`, so merging this starts nothing.
- Effort (added 2026-10-08 at head a6d2379): `--effort max` in `claude_args`, because judgment work (designated reads, adjudication, pull request and security reviews, audits) runs at `max` under the command center's effort mapping, and every job records its level and reason. The level has to be in `claude_args`: `--restricted` ignores the settings files that would otherwise carry it, and Opus 5.5 runs at `medium` on the API when a request leaves effort unset. The bounds test now requires the flag; removing it fails the test. A loopback dry run of the installed 2.1.295 client with this workflow's `claude_args` sent `output_config.effort: "max"` on every request.
- Fixes from the pre-cue read of the sibling #892 (2026-10-09, head d6f8909): the record names every behaviour of this new workflow: the trigger; the slug, owner, first-attempt and variable guards; the timeout and grants; the guard and binding exits; the head-as-data checkout and diff limits; the explicit action inputs `display_report` and `track_progress` (defaults in the pinned `action.yml`); and the 14-day usage artifact. The bounds step now uses an allow-list, fails on a start record without lists or on an empty result, and names each unmet bound. The review cut is character-safe and says when it cut. The step tests no longer skip without PyYAML. The 2026-10-03 cooldown waiver is cited.
- Fixes from this PR's own pre-cue toolkit read (J8 on d6f8909, head 8e64120):
  - the empty paths array is now expanded safely on bash before 4.4;
  - a success without an execution file now fails the job;
  - `display_report` and `track_progress` are asserted;
  - the artifact name and the action-pin sentence are corrected, and the parity receipt's dependency on #892 is stated;
  - `claude-pr-review.yml:review` is now in both grant tables of ci-least-privilege;
  - the record states that a public run's summary is public.
- Pre-cue toolkit read of this PR (J8 on d6f8909, 2026-10-09; `pr-test-analyzer` and `silent-failure-hunter`): every finding at confidence 80 or above, and its status at head 41c9c15.

  | Finding | Status at 41c9c15 | Fixing commit and evidence |
  | --- | --- | --- |
  | `pr-test-analyzer` 1 (90): the record names the wrong artifact | fixed | 8e64120: the record names `claude-pr-review-usage-<run id>-<attempt>`, the upload's name; 41c9c15 adds `test_the_usage_artifact_is_named_for_the_run_and_its_attempt` |
  | `silent-failure-hunter` 2 (88): the record and the workflow disagree on the artifact | fixed | the same two commits |
  | `pr-test-analyzer` 2 (85): `docs/github-automation.md` claims the harness audit's action pin | fixed | 8e64120: the bullet says the two pins differ until #892 lands |
  | `silent-failure-hunter` 1 (92): the same false claim (pin and federation variables) | fixed | 8e64120, as above; the four federation variables are the same |
  | `pr-test-analyzer` 3 (85): the headline evidence file is not in this PR | fixed | 8e64120 states that #892 adds it and lands first; 41c9c15 cites its `LR` entry field by field |
  | `silent-failure-hunter` 3 (80): "#897" and "two blocking findings" have no source in that receipt | fixed | #892 adds `target_pull_request`, `target_head`, `report_blocking_findings` and `report_sha256` to the `LR` entry (a marked 2026-10-09 amendment from the run's harness receipt); 41c9c15 cites them, calls the verdict the model's own, and states that the prompt and `claude_args` are unchanged from the run's workflow commit a6d2379 to that head (R5 later raises the two caps and the prompt's turn sentence; see R5 below) |
- R4, debug logging set as a repository secret or variable (2026-10-09; the cc-native-practice lane's L3 read against `anthropics/claude-code-action` v1.0.247 practice): the guard binds `(secrets.ACTIONS_STEP_DEBUG || vars.ACTIONS_STEP_DEBUG) == 'true'` and `(secrets.ACTIONS_RUNNER_DEBUG || vars.ACTIONS_RUNNER_DEBUG) == 'true'` as booleans and refuses either before any token is requested; a test pins both bindings, and six weakened copies each fail at least one test. This is hardening: the action step already pins `ACTIONS_STEP_DEBUG: 'false'`, which is what the action's full-output switch reads. zizmor's auditor persona reports `secrets-outside-env` on the two reads; CI runs the regular persona with `--no-ignores`, which reports nothing. The R4 lines are byte-identical in #892, #894, #895, #896 and #909; R5 leaves them unchanged here.
- The report notice at the cap (2026-10-09; the J8 micro read of #892): the publish step wrote with `jq -r`, which adds a newline, so a 60,000-byte report was announced as cut. It writes with `jq -j` now; a test publishes exactly 60,000 and 60,001 bytes, and a copy that writes with `jq -r` fails it. R5 adds the 59,999-byte case.
- From the J8 micro read of 41c9c15 (2026-10-09; 4 of 6 listed findings fixed, PTA-3 and SFH-3 partly because the receipt the record cites was not in this tree): this PR now carries a byte-identical copy of `local-parity-receipt.json` (sha256 `ba498b82…`, the file #892 adds), so the citation no longer depends on which PR lands first, and the docs bullet states the action pin and the federation variables instead of comparing with the harness audit on `main`.
- R5 (2026-10-09, head d1bcb6c, limited by the command center to correctness and test quality) adds a model-free step, `Remove symbolic links from the pull request head`, right after the head checkout: `set -euo pipefail; find pr-head -path pr-head/.git -prune -o -type l -exec rm -f {} +`. It raises the approved caps from measurement: `--max-turns` 12→30 and `--max-budget-usd` 3→5. The prompt's "You have at most 30 turns", the numbers step's bounds (1 to 30 assistant turns, at most $5, with matching messages), the header comment and `docs/github-automation.md` follow. The numbers step now counts each non-string tool-list entry as a forbidden tool named `non-string tool entry` instead of dropping it. The R4 guard lines, the deny list and the `github_token` input are kept exactly as at a1bc3bc. The test module goes from 33 to 41 tests. Eight are new: the symlink step's position and exact text; its behaviour on planted links to `/proc/self/environ`, `../.git/config` and a subdirectory, plus a re-run on a link-free head; a 59,999-byte report published whole; object, null and number tool entries refused and named; the prompt stating the `--max-turns` value; and the caps at 30 turns and $5 and just above. Changed tests pin the full `claude_args` list and `--settings` JSON exactly, assert the exact `if:` strings of the publish and numbers steps, run the guard with production's debug-off values (false/false/empty), and check that the 59,999-byte prefix before a cut character survives. All 28 weakened copies of the workflow fail at least one test with PyYAML present, as the hosted validate job has it; without PyYAML 15 of 28 do, because the pins sit in the PyYAML-only shape tests. The decision record gains a dated R5 section and names the R4 guard change in its what-changed paragraph. It corrects "Read denies for every `.git` directory" to "for the files under every `.git` directory". It restates the cap derivation and the expected cost per run, and withdraws the dry $0.70–$1.30 estimate, which was below the measured first-read cost. It adds the shared follow-up line, the shared `origin` URL sentence and two action sources. `manifests/evidence.json` is rehashed for the four changed files.
  - Sources of the R5 items: the command center's security read of a1bc3bc (the symbolic links, the closed flag list, the exact `if:` strings, the guard's debug-off values, the 59,999-byte prefix, and the stale guard sentence in the record); the command center's caps decision of 2026-10-09; the GPT designated read of #895, P2 (non-string tool entries, applied to all five W4 workflows); the cc-native-practice lane's third audit, P3-1 (the shared `origin` URL sentence).
  - Caps derivation: `LR`, a first read of #897 with this workflow's prompt and `claude_args` on the installed client, used all 12 of its 12 assistant turns and cost $2.23 ($1.87 by the client's estimate). So the need is above 12, and 2.5 times that floor gives 30 turns; twice $2.23 is about $4.5, rounded up to $5. Four delta re-reads (LR2 to LR5) used 5, 6, 7 and 6 turns and cost $1.75, $1.35, $1.54 and $0.73. A run is expected to cost about $1.9 to $2.3 for a first read and $0.7 to $1.8 for a delta re-read.
  - Filed as follow-ups before any enabling variable is set, not in R5: the security hardening from the 2026-10-09 read (debug-value widening, the tool-list shape apart from the non-string entry above, extra deny rules, token-source and guard-step assertions); the rest of F-02 (a non-empty tool list, checking the `tool_use` names); F-16 (the pins that run only with PyYAML).
  - Not run in R5: a fence smoke with a planted symbolic link, which needs a model run; any hosted run. On the first run, watch whether the action appends an `--mcp-config` that `--strict-mcp-config` still honours (the read's G3); the numbers step fails the run on any MCP server.
- Base commit: `aba02ec3` (merge base at a1bc3bc and at d1bcb6c, which builds on a1bc3bc; as read at a1bc3bc, main had moved to `ebe6c223` with #890, #900, #897 and #907; only `manifests/evidence.json` overlaps, and the one rebase happens at landing)
- Lane: `lane:foundation`
- Owned paths touched: `.github/workflows/claude-pr-review.yml` (new), `tests/test_claude_pr_review_workflow.py` (new), `tests/test_workflow_policy.py` (one inventory entry, one exemption), `tests/test_workflow_security_coverage.py` (coverage set and one zizmor test), `docs/decisions/2026-10-08-claude-actions-pr-review.md` (new), `docs/decisions/2026-10-04-ci-least-privilege.md` (dated addendum only), `docs/github-automation.md` (current-practice bullet), `evidence/artifacts/claude-actions-fence-smoke-20261008/receipt.json` (new), `evidence/artifacts/claude-actions-fence-smoke-20261008/local-parity-receipt.json` (new; byte-identical to the copy #892 adds), `manifests/evidence.json` (registration, last commit).

### R6 (2026-10-09), at bd10ba3

- **Cost bound:** the budget × 1.10 = $5.50, from four measured J8 runs on a $5 budget ($5.46, $5.33, $5.16, $5.0007),
  re-derived after the first three hosted runs.
  - A budget stop at or under it publishes the text written before the stop, with a line naming the stop.
  - Publishing is gated on `!cancelled() && steps.numbers.outcome == 'success'`, because the action fails its own step on
    any budget stop.
- **`--max-turns 30` removed:** the pinned action (`base-action/src/run-claude-sdk.ts:241-250` at 2dca132f, since
  6ef6450f) fails a success whose `num_turns` exceeds it.
  - On Claude Code 2.1.295, `num_turns` counts transcript messages: 57 for 12 requests under `--max-turns 12`. That is a
    code reading plus a local measurement, not a hosted run.
  - The numbers step's 30-assistant-turn bound and the prompt's "at most 30 turns" stay.
- **Job timeout:** 20 → 30 minutes. 30 turns at the measured pace (12 turns in 512,670 ms) take about 21.4 minutes before
  setup. #892 uses 30.
- **Settings pins:** both compare canonical JSON, so `1` no longer passes for `true` (J8 R5 micro, N1, confidence 70).
- **Record:**
  - the shared cost paragraph;
  - an upstream-issue draft (not filed);
  - the seven changed test expectations the J8 R5 micro found undeclared (U1–U7).
- **Local checks at bd10ba3:**
  - `validate.py` and the manifest check pass (10,906 files);
  - six modules: 262 tests, OK;
  - this module: 44 tests (16 skip without PyYAML);
  - actionlint and zizmor are clean;
  - 39 of 39 weakened copies fail a test with PyYAML.
- **Hosted checks at bd10ba3:** not read yet.

### SOTA sources

- anthropics/claude-code-action, v1.0.247 at `2dca132ff0e0c4094ce6048b422c6915a071210b` (<https://github.com/anthropics/claude-code-action/tree/2dca132ff0e0c4094ce6048b422c6915a071210b>): `docs/security.md` ("Do not check out an untrusted ref into the workspace root before this action"; base ref at the root and head ref in a subdirectory), `docs/setup.md` (federation inputs; "a static credential takes precedence and federation will not be used"), `action.yml` (inputs and the `execution_file` output), `base-action/src/parse-sdk-options.ts` (argument parsing with `shell-quote`, default setting sources, debug forcing full output), `src/github/operations/git-config.ts` (git authentication in the root checkout; lines 129-134 write the job token into the checkout's `origin` URL when `allowed_non_write_users` is unset) and, new in R5, `src/modes/agent/index.ts` (lines 52-61, where agent mode calls that configuration). R5 read both files at v1.0.246 (`38c80c1`); the command center's compare of v1.0.246 and v1.0.247 changes only the bundled Claude Code version.
- Claude Code 2.1.295, the version the pin installs, `claude --help`: `--restricted` ("removes the built-in tools that run commands or code ... and ignores user, project and local settings files ... Also confines the file tools to the working directories"), `--permission-prompts none` ("anything that would prompt is denied automatically"), `--tools`, `--settings`, `--strict-mcp-config`, `--max-budget-usd` and, for the R5 caps, `--max-turns`. The help does not say whether a symbolic link's target is resolved, which is why R5 removes the links before the model runs.
- GNU findutils `find` 4.10.0 as installed (`find --version`), new in R5: `find pr-head -path pr-head/.git -prune -o -type l -exec rm -f {} +`, measured by the step test on planted links and on a link-free head (exit 0, no `rm` started). The runner image's `find` was not measured.
- GitHub, Contexts reference, `runner.debug` (set only while debug logging is on, so the guard's `RUNNER_DEBUG_SIGNAL` is empty on a normal run), new in R5 as the basis of the guard tests' debug-off values; taken from the command center's read of a1bc3bc, not re-read in R5 (no network).
- GitHub, OpenID Connect reference, "Filtering for pull_request events" and "Filtering for a specific branch" (<https://docs.github.com/en/actions/reference/security/oidc>); GitHub Security Lab, Preventing pwn requests (<https://securitylab.github.com/research/github-actions-preventing-pwn-requests/>).
- Anthropic documentation read 2026-10-08: Workload identity federation (<https://platform.claude.com/docs/en/manage-claude/workload-identity-federation>); How Claude Code uses prompt caching (<https://code.claude.com/docs/en/prompt-caching>); Pricing (<https://platform.claude.com/docs/en/about-claude/pricing>).

### Evidence-class table

| Claim | Evidence class | Command / receipt |
| --- | --- | --- |
| With the fence flags (`--restricted`, `--tools`/`--allowedTools` Read,Glob,Grep, `--strict-mcp-config`, `--permission-prompts none`, `--settings` with hooks off, `claudeMdExcludes` and three deny rules; Haiku 5.5, 10 turns, $0.50, as the receipt lists; this workflow adds `--setting-sources user`, `--add-dir`, `--effort max`, four deny rules and `blockReadsOutsideWorkingDirectories`, a key in the 2.1.295 build that the receipt does not exercise) the session's tools are Glob, Grep and Read; reads outside the tree and of `.git/config` are denied; planted hooks, instruction files and a skill under `pr-head` have no effect | `native_proven` on the installed 2.1.295 client (not through the action, not on a GitHub runner) | `evidence/artifacts/claude-actions-fence-smoke-20261008/receipt.json`, three runs |
| The action's own parser turns this workflow's `claude_args` into those flags and that JSON | `local_integration` (at a1bc3bc only) | the parser's three helpers replayed with `shell-quote` 1.10.0 on the workflow text at a1bc3bc; not re-run at d1bcb6c (no network to install `shell-quote`). R5 changes two plain values (`12` to `30`, `3` to `5`), and `test_claude_has_three_read_tools_and_fixed_bounds` asserts the whole list, split per line with `shlex`, and the `--settings` JSON |
| The guard, binding, symbolic-link, diff, numbers and review steps behave as described | `synthetic` | `tests/test_claude_pr_review_workflow.py`, 41 tests at d1bcb6c (the 26 step tests run without PyYAML too; the 15 shape tests need it); each step's shell is taken from the workflow file and executed against a local stand-in for `gh`, a local git repository or a planted `pr-head` tree |
| Symbolic links under `pr-head` are removed before the model runs, `pr-head/.git` is left as the checkout wrote it, and a link-free head passes | `synthetic` | `test_symbolic_links_in_the_head_are_removed_and_nothing_else` (links to `/proc/self/environ`, `../.git/config` and `../..`; GNU find 4.10.0); no fence smoke with a planted link has run on the client |
| The tests detect a weakened workflow | `synthetic` | 20 mutated copies on 2026-10-08 (that day's bounds), 20 failed at least one test; on 2026-10-09 three turn-bound mutants each failed, and the new bounds have their own tests; at 41c9c15 three artifact-name mutants each pass the 30 earlier tests and fail the new name test; at d1bcb6c 28 mutated copies (the symbolic-link step, the flag and settings pins, the `if:` strings, the guard's debug-off values, the report cap, the non-string tool entries, the caps and the prompt) each fail at least one test with PyYAML 6.0.3 present, as the hosted validate job has it; without PyYAML 15 of 28 do (the other 13 are pins in the shape tests; follow-up F-16) |
| This workflow's own prompt and `claude_args` review a real pull request within bounds at `--effort max` on Opus 5.5 | `local_integration` (installed client, billed to a second key; not through the action) | `evidence/artifacts/claude-actions-fence-smoke-20261008/local-parity-receipt.json`, carried byte-identical in this PR and in #892 (sha256 `ba498b82…`), its `LR` entry: workflow commit a6d2379, whose prompt and `claude_args` are unchanged at d1bcb6c apart from the two caps R5 raised (`--max-turns` 12 to 30, `--max-budget-usd` 3 to 5) and the prompt's turn sentence; #897 at `target_head` 24821c5; 12 of 12 assistant turns; $1.868667 of the then $3 budget; tools Glob/Grep/Read; no MCP; `success`; `report_blocking_findings: 2`, the report's own verdict line, unchecked (`report_sha256` c49bee04, a 2026-10-09 amendment) |
| The caps, 30 assistant turns and $5, bound a runaway at about twice the measured need; a run is expected to cost about $1.9 to $2.3 for a first read and $0.7 to $1.8 for a delta re-read | `local_integration` (installed 2.1.295 client, billed to the second key; not through the action) | the `LR` entry above (12 of 12 turns, client estimate $1.868667); the $2.23 figure and the delta re-reads LR2 to LR5 (5, 6, 7 and 6 turns; $1.75, $1.35, $1.54 and $0.73) come from the command center's caps decision of 2026-10-09 and are not in this repository |
| In agent mode the action writes the job token into the root checkout's `origin` URL, so `persist-credentials: false` does not keep it out of `.git/config`; the deny rules keep the model from reading it | `source_review` for the code path; the denial of `.git/config` reads is the fence receipt's (first row) | `src/github/operations/git-config.ts:129-134` and `src/modes/agent/index.ts:52-61`, read at v1.0.246 (`38c80c1`; the command center's compare shows only the bundled client version changing at v1.0.247); the record's shared sentence also cites two 2026-10-09 probes on the installed 2.1.295, not reproduced here |
| The workflow is well formed and has no audit finding | `local_integration` | at d1bcb6c: `actionlint` 1.17.0 exit 0; `zizmor` 1.30.1 offline, regular (`--no-ignores`) and pedantic, no findings (2 suppressed in each: `secrets-outside-env` on the R4 secret reads, as at a1bc3bc) |
| A hosted run authenticates, stays in bounds and reads the cache | not established | no hosted run is part of this PR |

### Local commands run

```
$ python3 scripts/validate.py            # on the R5 tree, committed as d1bcb6c
{"components": 70, "hashed_files": 10906, "profiles": 4, "receipts": 239, "status": "passed"}   exit 0
$ python3 scripts/evidence_manifest.py --check
{"files": 10906, "status": "passed"}   exit 0
$ uv run --quiet --no-project --python 3.12 --with pyyaml==6.0.3 python -m unittest tests.test_claude_pr_review_workflow tests.test_workflow_policy tests.test_workflow_security_coverage tests.test_workflow_hardening tests.test_github_automation_practice tests.test_adoption_docs_consistency
Ran 259 tests ... OK (skipped=2)   exit 0
$ uv run --quiet --no-project --python 3.12 python -m unittest tests.test_claude_pr_review_workflow    # without PyYAML
Ran 41 tests ... OK (skipped=15)   exit 0
$ actionlint .github/workflows/claude-pr-review.yml
exit 0
$ zizmor --offline --no-config --persona=pedantic .github/workflows/claude-pr-review.yml
No findings to report. Good job! (2 suppressed)   exit 0
$ zizmor --offline --no-config --no-ignores --persona regular .github/workflows/claude-pr-review.yml
No findings to report. Good job! (2 suppressed)   exit 0
$ uv run --quiet --no-project --python 3.12 --with pyyaml==6.0.3 python -B <scratchpad>/r5_pr894_mutants.py . a1bc3bc    # outside the repository
28 mutants, 28 caught, survivors: []   exit 0
$ uv run --quiet --no-project --python 3.12 python -B <scratchpad>/r5_pr894_mutants.py . a1bc3bc    # without PyYAML
28 mutants, 15 caught   exit 1 (expected: the 13 survivors are pins in the PyYAML-only shape tests; follow-up F-16)
```

Hosted checks at d1bcb6c: not read yet.

### Decision record

`docs/decisions/2026-10-08-claude-actions-pr-review.md` (why a manual dispatch and no pull request trigger, how the head is handled, the model's fence and what was checked at runtime, cost of one run with the measured runs and the cap derivation, changed test contracts, alternatives, what would overturn it), with a dated R5 section: the symbolic-link step, the caps from measurement, the non-string tool entries, the closed flag list and the boundary tests, the mutants, and the follow-ups. `docs/decisions/2026-10-04-ci-least-privilege.md` gets a dated addendum naming the second job on the federation rule.

### Host evidence

Not applicable: no file under `evidence/hosts/` changes.

### Checklist

- [x] New/changed GitHub Actions are pinned to a full commit SHA with a
      version comment (no floating tags).
- [x] New/changed workflows declare top-level `permissions: {}` and grant each job
      only what it needs (`contents: read`, or an explicitly justified addition).
- [x] No secrets are printed, logged or committed; no new required secret was
      added without a documented owner.
- [x] No new paid hosting, subscription or billing surface was introduced.
- [x] Peer-owned untracked files and worktrees were preserved (not deleted,
      moved or overwritten).

Activation is a separate step after landing: an owner sets `CLAUDE_PR_REVIEW_ENABLED` to `true`, then dispatches the workflow with a pull request number and head commit. Runs authenticate by workload identity federation and spend from the Console organization the four `ANTHROPIC_*` repository variables name; no Anthropic key is stored in GitHub. Each run is capped at 30 assistant turns and a $5 client budget, the owner's per-run spend limit on that organization. This PR and #892 both edit the write-grant inventory and the evidence manifest, so the second to land rebases once.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant