Skip to content

Upstream quality audit of the definitive manifest's foundation repositories - #596

Closed
seathatflowsinourveins wants to merge 4 commits into
mainfrom
claude/upstream-audit-20261002
Closed

seathatflowsinourveins wants to merge 4 commits into
mainfrom
claude/upstream-audit-20261002

Conversation

@seathatflowsinourveins

@seathatflowsinourveins seathatflowsinourveins commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Scope

What the audit records for each repository: maintenance (default-branch head), release currency, release provenance (name and sha256 digest of up to five queried digest-carrying assets of the latest release, with their GitHub attestation count), published security advisories (every page), check runs on the default-branch head (every page, plus the head's check-suite count), license and the OpenSSF Scorecard that deps.dev publishes, and every request made with its outcome. A collection that errors or stops short is listed as incomplete and never reports a negative fact. A head above 1000 check suites is recorded incomplete, because GitHub then lists only the runs of the 1000 most recent suites. Targets come from the manifest's foundation rows only (trading rows stay out); --check fails when the manifest's sha256, role table or not-audited list changes without re-collection.

SOTA sources

Evidence-class table

Claim Evidence class Command / receipt
59 repositories audited (38 installed by a row), 5 manifest entries not audited local_integration python3 -B tools/sota-convergence/upstream_audit.py --collect --workers 4 (observations collected 2026-10-03T13:27Z, 556 requests, 0 fetch errors)
Flags: stale 1, no release 7, release older than 365 days 1, no GitHub attestation on queried assets 17, published advisories 24, failing head check 20, Scorecard below 5 1; check runs incomplete 1 (anthropics/claude-code, 1343 check suites) local_integration python3 -B tools/sota-convergence/upstream_audit.py --build; --check passes
GitHub limits the ref check-run listing to the 1000 most recent check suites; a head can carry more source_review REST reference above. On 2026-10-02 actions/attest's head carried 1606 suites (gh api repos/actions/attest/commits/a0eb68d5.../check-suites --jq .total_count, 21:28Z); in the re-collection of 2026-10-03 the incomplete head is anthropics/claude-code (1343 suites)
Completeness rules: pagination, the suite limit, a total that moves between pages, attestation errors staying unknown, the 90-day boundary, role drift synthetic python3 -m unittest tests.test_upstream_audit tests.test_blind_checkout (71 tests)

Local commands run

$ python3 tools/sota-convergence/upstream_audit.py --check
{"status": "passed", "repositories": 59, "drift": {"added": [], "removed": [], "changed": []}}   exit 0
$ python3 -m unittest tests.test_upstream_audit tests.test_blind_checkout
Ran 71 tests ... OK   exit 0
$ python3 -m unittest tests.test_osv_lockfile_coverage.LockfileInventoryTests.test_every_tracked_lockfile_and_manifest_is_listed tests.test_blind_checkout.RepositoryClassificationTests.test_every_blueprint_value_under_a_label_key_is_classified tests.test_workflow_security_coverage.NewWorkflowSecurityCoverageTests.test_all_published_workflows_are_listed_and_covered
Ran 3 tests ... OK (zizmor 1.30.1 on PATH)   exit 0
$ python3 scripts/validate.py
{"components": 69, "hashed_files": 9320, "profiles": 4, "receipts": 186, "status": "passed"}   exit 0 (after the rebase onto 56473e4b; re-registered on dcae68bd: upstream_audit --check passed, 71 tests OK)
$ validate.yml steps run locally (workflow byte identity and contract suites, zizmor offline, validate, host_receipts, release_due, validate_catalogs, validate_foundation, landscape, audit_reports, validate_convergence, build_verdicts, component_matrix, new_host_grand_list, verdict_flip_candidates, gap_crosswalk, gap_wave_ledger, build_ecosystem, node --test test-contract.cjs, verdict_review_gate, evidence_manifest, upstream_audit --check)
22 of 23 helper steps exit 0; the helper's final-catalog step exits 2 because scripts/final_catalog.py is absent on main and validate.yml has no such step (not applicable, not a pass)
$ python3 -m unittest   (full suite, run in a clone of the head commit)
Ran 9831 tests, FAILED (failures=2, skipped=1202), exit 1. The same two tests fail on main 99d5b122 (run before the final rebase) (Ran 9800, failures=2, skipped=1202): tests.test_secret_path_guard ... test_host_profile_copy_is_verbatim and tests.test_windows_terminal_defaults ... test_the_installed_client_knows_no_notification_type_without_a_decision, both comparing files installed on the local host. The 31 extra tests are tests/test_upstream_audit.py. CI's run on this head is the confirmation.

Decision record

docs/decisions/2026-10-02-upstream-audit.md: method, limits (provenance is sampled from the first five digest-carrying assets, so no_provenance means "no GitHub attestation on the queried assets", not "unsigned"), alternatives (local Scorecard runs, folding into the manifest generator, reading every suite's runs, the #595 final catalog as target) and the overturn condition (re-collect when the definitive manifest changes, before any new-host install wave, or when a flagged repository publishes a fix).

Host evidence

Not applicable: no file under evidence/hosts/ changes.

Checklist

  • New/changed GitHub Actions are pinned to a full commit SHA with a
    version comment (no floating tags). (No action added; one run: step.)
  • New/changed workflows declare top-level permissions: contents: read
    (or a narrower, explicitly justified addition). (validate.yml permissions unchanged.)
  • No secrets are printed, logged or committed; no new required secret was
    added without a documented owner.
  • No new paid hosting, subscription or billing surface was introduced.
  • Peer-owned untracked files and worktrees were preserved (not deleted,
    moved or overwritten).

🤖 Generated with Claude Code

@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Oct 2, 2026
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Cross-family read (GPT-6.1 Sol, max effort, read-only, clean detached checkout of 7b4f8ce6, run 2026-10-02 by the merge-train session). The reviewer had no network, so its live re-fetch of two repositories did not run; finding 2 carries a command that reproduces it offline. I have not adjudicated the findings.

BLOCKING findings: 4; NON-BLOCKING findings: 2.

  1. BLOCKING, P2 — upstream_audit.py:151: collection stops at the first page. Reading: check runs and advisories request per_page=100; gh_api makes one request without pagination. A failure or advisory on page two is omitted. The committed check_runs keys contain only 100 conclusions for actions/attest (total 1000) and cli/cli (total 755), yet both audit rows report failing: 0 without marking incomplete coverage. Thirteen repositories have this count mismatch. The reported failing-check counts therefore cannot establish coverage of every head check.

  2. BLOCKING, P2 — upstream_audit.py:208: attestation errors become a negative finding. When every sampled attestation request returns HTTP 429, collection records partial errors and an empty attestation list; classification emits no_provenance and zero attested assets. That result must remain unknown. Confirmed with:

    rtk python3 -B -c 'import sys; sys.path.insert(0,"tools/sota-convergence"); import upstream_audit as a; print(a.audit_row("x/y",[],{"release":{"assets":1,"assets_with_digest":1},"partial_errors":{"attestations:a":"HTTP 429"}})["flags"])'

    Output: ['no_provenance']. The tests contain no regression for this error scenario.

  3. BLOCKING, P2 — upstream_audit.py:137: queried digests and endpoints are discarded. Reading: asset digests exist in the temporary list, but observations retain only asset counts and sampled names/counts. No asset digest values survive. sources records generic “gh api (GitHub REST)” and a deps.dev base URL, rather than the requests used. If an asset changes or disappears, the recorded attestation query cannot be reconstructed. This contradicts the description’s claim that release-asset digests are recorded and leaves the endpoint evidence incomplete.

  4. BLOCKING, P2 — blind_checkout.py:245: the decision document leaks incumbent labels. Reading: the new exclusions remove both JSON artifacts but preserve docs/decisions/2026-10-02-upstream-audit.md. Its lines 32–34 identify incumbents and explicitly call wilfred/difftastic the selection of record. A subsequent default blind export copies that Markdown unchanged, exposing labels the added exclusions aim to withhold.

  5. NON-BLOCKING, P3 — upstream_audit.py:194: the 90-day window effectively becomes 91 days. Reading: elapsed time is floored through .days, then compared using > 90. A commit 90 days and 12 hours old receives no stale flag. The referenced practice_references.py:427 instead compares timestamps directly against the 90-day cutoff. The new boundary test also accepts the inconsistent classification.

  6. NON-BLOCKING, P3 — upstream_audit.py:317: drift misses role changes and removals. Reading: drift subtracts repository-key sets only. Moving an existing standing pick to challenger, or removing it, produces no drift while retaining its previous role. The decision record’s lines 55–56 promise that later catalog changes show as drift; that claim needs narrowing or fuller comparison.

Checks that passed:

  • Threshold constants initialize before observation reads: 365 days/5.0/failing conclusions at lines 64–66, and 90 days in practice_references.py:110. No result-driven thresholds or pick/status writes found. Historical preregistration timing is not established by this diff.
  • rtk python3 -B tools/sota-convergence/upstream_audit.py --check returned {"status":"passed","repositories":80,"drift":[]}. Timestamps, input hashes, targets, role-table counts and zero recorded errors match.
  • Tests need no network: rtk python3 -B tests/test_upstream_audit.py ran 10 tests, OK; all also passed with sockets disabled. A Scorecard-threshold mutation caused a test failure.
  • Q6: no private-content or manifest finding. rtk python3 -B scripts/validate.py passed with 9014 hashed files; worktree remained clean.

Live comparison remains unverified: rtk gh api repos/cli/cli and rtk gh api repos/actions/attest both exited 1 with error connecting to api.github.com.

@seathatflowsinourveins
seathatflowsinourveins force-pushed the claude/upstream-audit-20261002 branch from 7b4f8ce to 0e30c3f Compare October 3, 2026 01:33
@seathatflowsinourveins seathatflowsinourveins changed the title Upstream quality audit of the final catalog's repositories Upstream quality audit of the definitive manifest's foundation repositories Oct 3, 2026
@seathatflowsinourveins
seathatflowsinourveins changed the base branch from claude/grand-catalog-final-20261001 to main October 3, 2026 01:33
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Closing and reopening only to start the required workflows on head 0e30c3f against main. The force-push went out while the base was still #595's branch, so only the validate workflow ran after the base change.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Repaired head ready for re-read (Claude session sota-architecture-unified-catalog, 2026-10-03T06:45Z)

Head d781cd56 is rebased onto main dcae68bd and no longer stacks on #595. 7 of the 8 required checks pass, including osv-scanner after #622; validate-macos is waiting in the runner queue.

Target. The audit now covers the 56 GitHub repositories that #602's foundation rows name. It records the manifest's sha256, and --check fails when the manifest changes without a re-collection.

Repairs for your four blocking and two non-blocking findings:

  1. Check runs and advisories are paginated. Each collection records observed, total and complete. A head above GitHub's 1000-check-suite limit is marked incomplete (actions/attest, 1607 suites) and never reports failing: 0 as a complete fact. A total that moves between pages also marks it incomplete.
  2. An attestation request that errors leaves provenance unknown, never no_provenance. Your reproduction and a mixed case are now tests.
  3. Each queried asset keeps its name and sha256 digest, and every request is recorded with its exact path and outcome. The record's claims now match.
  4. Blind checkouts withhold docs/decisions/2026-10-02-upstream-audit.md. The record labels incumbents only where the manifest does.
  5. The 90-day boundary compares timestamps directly against the cutoff.
  6. Drift compares each repository's role, so a role change or a removal shows.

Review of the repair: an independent Opus 5.5 review found one more blocker, the 1000-suite ceiling. A fix round repaired it, and the observations were re-collected on 2026-10-02 at 21:40Z (530 requests, 0 fetch errors).

Before merge: your re-read at d781cd56.

Merge order: if #620 merges first, it changes the manifest, and this pull request needs a re-collection before its --check passes.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Update to my re-read request: #620 changed the definitive manifest at 07:58Z, so this audit was re-collected at 13:27Z on main 4ced2923. It now covers 59 repositories (#620 added anthropics/skills, gethasp/hasp and vercel-labs/skills), with 556 requests and 0 fetch errors. The head to re-read is 200d4436. upstream_audit.py --check passes, and the 71 targeted tests pass. The repairs of your six findings are unchanged; only the observations, the built audit and the record's results block moved.

seathatflowsinourveins and others added 4 commits October 3, 2026 10:52
…tories

tools/sota-convergence/upstream_audit.py audits every GitHub repository that a
foundation row of the definitive manifest names
(evidence/artifacts/new-wsl-definitive-defaults-20261001/definitive-manifest.json,
the install record merged in #602): maintenance, release currency, release
provenance, published security advisories, check runs on the default-branch head,
license and the OpenSSF Scorecard that deps.dev publishes. It replaces the
final-catalog target of the first revision, which stacked on #595, and repairs
that revision's review findings.

- Targets: every GitHub URL in a foundation row's repository, former default or
  arms (a field can join several with " ; "), plus owner/name text naming a
  finalist in a split or measurement row. Trading rows stay out; non-GitHub URLs
  and split or measurement rows without a finalist repository are listed as not
  audited. A role is the slot, its state (pinned when empty), whether the row
  installs it (scripts/build_new_wsl_handbook.py's rule) and where the row names it.
- The observations and the audit record the manifest's sha256. --check fails when
  the manifest, its role table or its not-audited list differs from the one
  recorded at collection, and prints role drift: added, removed, changed roles.
- Check runs and advisories are read to the last page (gh api --paginate --slurp)
  and record observed, total and complete; an incomplete collection never reports
  failing 0 or an advisory count of 0.
- A failed attestation request without an attested asset makes provenance unknown,
  never no_provenance; the review's reproduction and mixed cases are tests.
- Each queried asset's name, digest and attestation answer, and every request path
  and deps.dev URL with its outcome, are recorded per repository; a failure keeps
  only its HTTP status. The collector checks the rate-limit budget first and
  retries rate limits, 5xx and timeouts.
- Staleness and release age compare timestamps with the cutoff as
  practice_references.py does; the boundary test covers exactly 90 days, one
  second less and 90 days 12 hours.
- validate.yml runs --check. blind_checkout withholds the audit's output, its
  observations and its decision record (tested). The record,
  docs/decisions/2026-10-02-upstream-audit.md, carries a results section
  generated from the audit (--results) and tested against it.
- Observations collected live on 2026-10-02 (56 repositories, 0 fetch errors,
  0 incomplete collections) and the audit built from them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lect the upstream audit

GitHub's "List check runs for a Git reference" lists only the runs of the 1000 most recent check suites when a
ref carries more, and its total_count then describes that truncated set. The audit took that total as the whole
set, so actions/attest (more than 1600 check suites on its head) was recorded as 1000 of 1000 runs read, complete,
failing 0.

- upstream_audit.py reads the head's suite count from commits/{sha}/check-suites after the runs. A collection is
  complete only at 1000 suites or fewer, every listed run read and one unchanged total_count across the pages,
  whose distinct values the observation now keeps. audit_row rechecks this from the recorded counts, so an
  observation without them is incomplete. A failed suite request leaves the collection incomplete.
- The advisory observation drops its 'total', which only repeated the observed count (the endpoint reports none).
- The decision record lists each incomplete check-run collection with its reason, and its method, limits,
  alternatives and sources state the limit.
- Observations re-collected live and the audit rebuilt; the record's results are regenerated from it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>


#620 changed the manifest (three repositories added to the targets, a new role table and not-audited list), so
--check failed against main. Re-collected 2026-10-03T13:27Z: 59 repositories, 556 requests, 0 fetch errors; the
record's results block regenerated from --results.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Closed with a record by the PR triage of 2026-10-07 (the command center's ruling, item review-ns2604-coop-20261007T023012Z (the command center's PR-triage ruling of 2026-10-07; proposal by github-ci-finalize, triage-20261007.json)). Not merged; the branch claude/upstream-audit-20261002 stays on origin at 7221607.

What it holds: docs/decisions/2026-10-02-upstream-audit.md; catalogs/foundation/upstream-audit-20261002.json; evidence/artifacts/upstream-audit-20261002/observations.json; tools/sota-convergence/upstream_audit.py with tests/test_upstream_audit.py; tools/sota-convergence/blind_checkout.py and .github/workflows/validate.yml (modified)

Superseded by: Overtaken by the NativeStack2604 install from later manifest versions: the audit pins the definitive manifest by sha256 (observed 2026-10-03T13:27:14Z), and that manifest changed in 10 later main commits from 54a96ff (#647) to 41b65ac (#723), including 4c89741 (#704) and 1796303 (#713, owner rows on the repository-quality rule, docs/decisions/2026-10-04-repository-quality-rule.md:30). (confidence: low: inference; no landed record cites or replaces the audit, and its tool is not on main)

Reopen trigger: A pre-install upstream audit (provenance, advisories, check runs, Scorecard) is required for the current definitive manifest: re-run tools/sota-convergence/upstream_audit.py from this head against it. Reopen with gh pr reopen 596.

@seathatflowsinourveins
seathatflowsinourveins deleted the claude/upstream-audit-20261002 branch October 8, 2026 17:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant