Repository navigation
Record the macOS full-suite coverage decision, the CI measurements and five anti-patterns - #632
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 13e26ea50e
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Hosted-check note for the merger: the required |
13e26ea to
8e12eed
Compare
…d five anti-patterns The 2026-10-03 GitHub CI review measured the required test jobs: both run the whole suite (about 9,500 tests) serially, Linux validate in 1,727-1,828 s against a 2,400 s limit and macOS validate-macos in 1,932-2,394 s, and 30 of 1,025 pull-request head SHAs (2.9%) failed only on macOS in 8 days. The decision record keeps full macOS coverage on every pull request (O4), pre-registers a fallback that gates only the full-suite step (O2, not adopted, with a prospective shadow phase, a frozen escape bound and a derived INERT list), rejects gating the whole job (O1, forbidden by tests/test_workflow_hardening.py:1247-1255) and running the suite only after merge (O3), and lists reductions R1 to R6 for their owners. The measurements are kept as a durable receipt because Actions run ids expire with the new retention period; the itemized failing runs carry their head SHAs. Five proven mistakes join the anti-pattern log. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Answer the review threads on the receipt with published evidence: - Publish the compiled measurement input byte-identical (evidence/artifacts/github-ci-measurements-20261003/input.json, sha256 d01f9b63...) and the read-only GET calls of 2026-10-03 that check the receipt (later-reads-20261003.json: 218 calls with argument vectors, times, exit codes and payload hashes; payloads withheld because they carry commit author emails). - Add a queries section: request forms reconstructed from the measurement plan, not a transcript, and what was not retained. The receipt is a recorded summary that can be re-derived only while the run records exist; durability wording in scope and limitations is replaced. - Caches: record 130 (likely the usage endpoint) and 131 (the by-prefix sum) with the byte reconciliation; the difference stays unreconciled, and two later paired reads show the same one-cache gap. - Failure anatomy: the 80 of 88 Validate runs are the failed runs no later same-SHA run turned green (reproduced exactly); the 8 omitted are sota-sources-only failures (21 of 94 instances over all 88). Other workflows: 10 of 76 classified, 66 not. - Self-audit fixes: the 17-entry failing-test key and limits string, the composition of the 11 non-pull-request cancelled runs, the 96 and 87 denominators, the sampled failed-job seconds and the suite-shape definitions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- O2's prospective phase becomes a dry run: the required validate-macos job keeps running the full suite on every pull request, the gate only reports what it would skip, E_G is read from the required job on the would-skip pull requests, and skipping starts only after the preregistered thresholds pass in a later reviewed change. The thresholds are unchanged. - R3 and R5 state the failure anatomy's coverage (80 classified of Validate's 88 failed pull-request runs, all 75 of Adoption's, 10 of 76 other-workflow failures) and what the 8 omitted runs change (sota-sources 21 of 94 over all 88); R5 names the steps behind 18 and 15. - The measured findings and evidence class say the receipt is a recorded summary that can be re-derived only while the run records exist, and point to the published input and the later reads. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the data Second review of the receipt and the record (R1-R5): - R1: the reclassification of all 88 failed Validate runs forces each of the 8 runs the anatomy left out to have failed only in sota-sources (94 - 86 = 8 instances, 21 - 13 = 8 of them sota-sources), but any 8 of the 19 sota-sources-only failures reproduce the counts. The 8 that a later same-SHA run turned green are consistent with the input's tally of 8 of 96, not established as the 8 left out: 96 covers all events and the 6 push and 2 workflow_dispatch failures were not checked. The record says once that R3 and R5 do not depend on which 8. The receipt's coverage keys become selection and runs_with_a_later_success. - R2: the byte total equals the by-prefix sum to 0.1 MiB but does not indicate which cache count is right; the later listing's total_count of 129 is stated. - R3: the measuring agent's request forms give the path only (client flags not retained); the GET statement is scoped to the input and a GH_DEBUG check that was not retained; the suite-shape figures are not covered by the plan; the queue-seconds definition is marked inferred; the two-endpoint cross-check is relayed as the input's statement. - R5: largest_file reads 632 'def test' occurrences (624 test definitions) in 8,441 lines, listed in source.changes_from_input. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the same 8 in the anatomy selection text Delta review findings on the CI measurements receipt: the later-success creation time is in no published artifact, so the sentence now points at the run ids in the later-reads artifact; the reclassification of all 88 runs keeps every other by_step cell equal but policy_gate rises by the same 8 because it counts sota-sources. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…e anti-pattern log (hot-file protocol; last commit) One registration on main's manifest at 4ced292: the receipt, the decision record, input.json, later-reads-20261003.json and docs/harness-defaults.md (five rows). It is the branch's last commit, so a later rebase takes main's manifest and registers again (docs/lanes.md). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
8e12eed to
85e9bc6
Compare
…ify the six cache calls' batch timestamps in the receipt (review 632 P1, P2) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…egistry plus the owned rows) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fore the final hot-file commit Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… (hot-file protocol: every hot-file edit in the last commit) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…egistry plus the owned rows) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fore the final hot-file commit Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… (hot-file protocol: every hot-file edit in the last commit) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Claude session native-agent-stack-5f: landing at head Observed main Required checks at this head: 7 pass 1 skipping . Unresolved review threads: 0. |
|
Claude session native-agent-stack-5f: post-merge observation. Landed as |
…Sol round 642-merge-r3) Resolutions: - docs/harness-defaults.md keeps both sets of anti-pattern rows (this PR's, then main's #632). - The new-WSL handbook and its receipt are regenerated by scripts/build_new_wsl_handbook.py on the merged sources. Four wording P2s from the Claude read are fixed: the dated amendment line, the HUD block citing main's 2026-10-04 receipt, commits/<tag>, and the path count. Acceptance: validate 0, evidence_manifest --check 0, handbook/matrix/grand-list/build_ecosystem --check 0, 132 tests (120 OK, 12 skipped), git diff --check 0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scope
4ced2923063db6a6dcafa9f25af5ee05a4153c75(rebased ontomainon 2026-10-03 after Track GPT runtime workers, SDKs, the gateway and pi in the daily catalog-freshness report #634 and AGENTS.md: drop "(lands with unit F3)" now that the skill lifecycle guide is on main #636; the file and line citations in the record were taken at56473e4b840f0e6940c031801d866e7e9bf29baf, the commit the work was cut from)lane:foundationdocs/decisions/2026-10-03-macos-full-suite-coverage.md,evidence/receipts/github-ci-measurements-20261003.json,evidence/artifacts/github-ci-measurements-20261003/(input.json, the compiled measurement input, byte-identical;later-reads-20261003.json, the 2026-10-03 GET calls with argument vectors, times, exit codes and payload hashes),docs/harness-defaults.md(five rows at the top of the anti-pattern table),manifests/evidence.json(registration only, the branch's last commit, per the hot-file protocol indocs/lanes.md)SOTA sources
56473e4b(the record's citation commit; at the base commit the sametests/test_workflow_hardening.pylines are 1244-1257):tests/test_workflow_hardening.py:1242-1255,docs/acceptance-evidence-policy.md:42-53and:57-60,.github/workflows/validate.yml,.github/workflows/adoption-bootstrap.yml.Evidence-class table
gh, 2026-10-03), kept as a recorded summary (the published input), not as returned payloads; nearest template classsource_reviewevidence/receipts/github-ci-measurements-20261003.json;evidence/artifacts/github-ci-measurements-20261003/input.jsonsota-sources, so R3 and R5 do not depend on which 8; that they are the 8 a later same-SHA run turned green is consistent with the input's tally of 8 later-success runs of 96, all events, but not established: any 8 of the 19sota-sources-only failures reproduce the counts, and the 8 non-pull-request failures were not checked), the 712 cancelled runs by workflow and event, both cache endpoints and the retention settingevidence/artifacts/github-ci-measurements-20261003/later-reads-20261003.json; receiptsource.later_readsdata.caches_and_artifacts.cacheslocal_integrationLocal commands run
At
55425797(this branch's head, the registration commit):An uncommitted arithmetic and consistency self-audit of the receipt, the published input, the later reads and the record, extended for the second review, ran 243 checks at
55425797, with 0 failed. 25 of them fail on the previous head69bd8d52and pass now: the 24 checks of the new wording and the check that the receipt's data differ from the input exactly wheresource.changes_from_inputsays. Five more check the data behind the new wording (19 runs that failed only insota-sources, 1, 14 and 4 of them created on 2026-09-25, 09-27 and 09-28; the 8 later-success runs among them; at least one failing instance per run; the 94 - 86 and 21 - 13 differences) and the unchanged coordinator request forms.Decision record
docs/decisions/2026-10-03-macos-full-suite-coverage.md(alternatives, the O2 adoption rule fixed in advance with a dry-run phase that keeps the required job's full suite, overturn conditions).Host evidence
Not applicable: no files under
evidence/hosts/change.Checklist
permissions: contents: read(no workflow changes in this PR).Notes for the merger:
manifests/evidence.jsonis the shared hot file (rebase and re-register perdocs/lanes.md). This branch's five 2026-10-03 anti-pattern rows sit at the top of the table indocs/harness-defaults.md, above the 2026-10-02 row from #634. If #635 or another open PR also adds 2026-10-03 rows at the top of the same table (the pushed head7ce51c0aof #635 did not changedocs/harness-defaults.mdwhen checked on 2026-10-03), whichever merges second keeps both sets, newest first, and re-registers. If this branch merges second: takemain'smanifests/evidence.json, registerdocs/harness-defaults.md,docs/decisions/2026-10-03-macos-full-suite-coverage.md,evidence/receipts/github-ci-measurements-20261003.jsonand the two files underevidence/artifacts/github-ci-measurements-20261003/again, then run the three--writecommands. Open PRs #641 and #648 also change rows of this table. This branch is based onmainat4ced2923(after #634 and #636);git merge-tree --write-tree --name-only origin/main HEADwas clean when it was pushed. The record deliberately names no private repository and no dollar figure.🤖 Generated with Claude Code
Custody refresh, 2026-10-04
docs/decisions/2026-10-03-macos-ci-scope.md) replaced this record's coverage policy. Commit22e48ee0fadds a dated note under the title and a historical marker on the Decision, Alternatives, Preregistered checks and Overturn sections; the current macOS coverage policy is macOS CI scope: validate-macos stays required, running the full suite for Mac-relevant changes, only changed test modules for test-only changes, and skipping otherwise (decision record, receipt, tests) #677's. The Measured findings, the receipt and its measurements stand.source.later_reads.whatnow states that the six cache calls carry their enclosing batch's start and end times, not per-command times, as the artifact's note says.c4900d7f3merges main14048b84. The anti-pattern table conflict was an insertion on both sides, so both sets of rows are kept. The registry is in the last commit (b9e43d39e).85e9bc65returned FINDINGS (one P1, one P2). The delta read atb9e43d39returned ACCEPT with both resolved.