Repository navigation
Track GPT runtime workers, SDKs, the gateway and pi in the daily catalog-freshness report - #634
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
80b5537 to
2d28e04
Compare
|
Cross-family read at The reviewer was Findings:
Round 5 repairs all four:
The same reviewer then rereads the repaired head before this PR is merged. |
2d28e04 to
f6a2f68
Compare
|
Cross-family reread at
The reviewer found no regression. It noted that after the rebase, #639 makes scheduled proposals Monday-only, which does not disturb this PR's build, report or upload steps. It ran 30 read-only tests and checked the seven re-registered hashes. |
…ness report The daily catalog-freshness report covered the foundation stack and the trading pins only. The GPT-route runtime workers, agent SDKs, the GPT gateway, the evaluation harnesses and coding agents such as pi were not compared with upstream at the pins their runtime records install, so the foundation the trading north star runs on could fall behind without notice. extract_layers.py gains RUNTIME_PIN_SOURCES, 12 pins read at extraction time from the new-WSL install plan rows by slot (a composite row by part), the OpenHands runtime-worker recipe pin record and the native SDK constraints (PEP 503 name match), and RUNTIME_WATCH_SOURCES, 7 watch-only upstreams that a file on main names but no record on main pins (pi, oh-my-pi, the OpenAI Agents SDKs, Crawl4AI, Deep Agents and codex-action). Both resolve into runtime-pins.json. A malformed declaration raises. A record that moved or changed shape yields an entry with "pin": null and an error built only from the declared path and locator, so the daily job keeps its foundation and trading report. github_freshness.py fetches these repositories. build_manifest.py --runtime-freshness-out writes the report-only runtime-freshness.json; the manifest and the trading sidecar stay byte-identical. freshness_propose.py renders it as a separate drift.md table that never sets drift-status.txt and that the propose job never reads. The workflow writes and uploads the sidecar and prints its counts. These rows stay out of TRADING_PIN_SOURCES, whose pins must not repeat a selected card's repository (openai/codex, OmniRoute, inspect_ai and deer-flow are cards) and must sit in a trading taxonomy layer. tests/test_catalog_freshness_runtime.py covers extraction against the checked-in records, the resolver on synthetic records, the fetch, build and report stages and the workflow wiring, with no network. scripts/validate.py reports SHA-256 and byte-count mismatches for the six registered files edited here; they are re-registered in the branch's last commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… runtime tests Resolves the round-2 review findings on 6bc1199 (none blocking or major). F1: build_manifest.py catches LeakDetected around the runtime sidecar's gate only. A runtime upstream whose own data trips the gate (a third-party tag containing APCA, say) no longer fails the build step and so loses the workflow's diff step and drift report: the sidecar keeps its schema and keys with no entries, every count 0 and the fixed "gate_error": "leak_gate_tripped" (never the matched text or the exception message), the step prints {"runtime_freshness": {"gate_error": "leak_gate_tripped"}} and exits 0. The manifest's and the trading sidecar's leak checks stay fatal. freshness_propose.py renders that document as the runtime heading and one withheld line with no table, and returns runtime_unresolved ["leak_gate_tripped"], so the workflow's counts line shows it. F2: the install-plan test asserts each entry's declared slot and, for the two research-harnesses parts, the part_repository slug at its position. The codex and codex-sdk-and-codex-exec-app-server rows carry identical values, so a swapped slot passed before. F3: watch-only now means no install or runtime record on main pins the upstream (a catalog card may record an evaluated version). The alternative openai-agents-sdk card records v0.22.3 for openai/openai-agents-python, which no table compares. The comments, the build docstring, the README and the drift.md intro say so, and that the stack.json and selected (default/conditional) card pins stay in the other tables. F4: the synthetic constraints give openai-codex-cli-bin 9.9.9, so the PEP 503 test proves the openai_codex line matched. F5: the README says an unresolved row carries its upstream and dormancy only when it has a repository; one without has an empty upstream and the not_fetched dormancy. F6: a test runs extract_layers.main with a call-through spy on resolve_runtime_pins and checks that reserved_ids holds hftbacktest (a trading pin), nautilustrader (a trading card) and codex (a foundation component). F7: an unresolved row without a repository is no longer also listed as unfetched; the unresolved line reports it. A row that kept a declared repository is still listed when it was not fetched. Mutation runs in a scratch copy fail each new assertion on its target defect; the base test file passed the F2 slot-swap and F4 prefix-match mutants. scripts/validate.py reports SHA-256 and byte-count mismatches for the same six registered files as 6bc1199; they are re-registered in the branch's last commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…orrection - docs/decisions/2026-10-02-gpt-runtime-tracking.md: the gap by row kind, the tracked entries and their pin records, ownership (the Codex maintenance task owns upstream detection; this table is a report) and the relation to open #633's runtime-job catalog, the search-first sweep (wf_21fc37c5-123: extend the existing job; updatecli, Renovate and nvchecker trialed on copies of the records), the first local run against live GitHub metadata, verification, limits and follow-ups (tag-only rows, registry versions, advisories through OSV-Scanner). - docs/harness-defaults.md: anti-pattern row for stating what a catalog job tracks from where a repository's name appears, inserted at the top of the log table (open #619 appends at its end). - tests/test_catalog_freshness_runtime.py: the leak-gate test also captures stderr and checks that neither stream carries the matched text or the exception message; the reserved-id test checks that each id belongs to exactly one report kind. Both review nits; the stderr check fails when the caught exception is printed (mutation run). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Inspect AI publishes no GitHub release and its first tag in name order is
release/2025-11-28, so new-wsl:inspect-ai was not compared although its newest
version tag is 0.3.276 against the pin 0.3.273. watch:codex-action (tags only)
and watch:deepagents (a monorepo whose latest release can be another package's)
relied on name order or on whichever package released last.
- extract_layers.py: an optional "tags" declaration {prefix, pattern} on a
runtime pin or watch source, checked at declaration time (a literal prefix, a
pattern anchored with ^ and $ that has exactly one capture group), declared
on new-wsl:inspect-ai (every tag), watch:codex-action (v) and
watch:deepagents (deepagents==); each entry carries it or null.
- github_freshness.py: one "gh api repos/{slug}/git/matching-refs/tags/{prefix}
--paginate" call per declared prefix, tag names stored under matching_tags;
a failure is a partial error that never aborts the batch, and a record
without a list for a declared prefix stays pending.
- build_manifest.py: build_runtime_freshness takes the highest version among
the names the pattern fully matches (the capture as an integer tuple, never
name order) as upstream.latest before the pin is compared, with latest_source
and matching_tag_count; otherwise tag_pattern_unmatched or
tag_pattern_unfetched.
- freshness_propose.py: marks a matching-tag latest "(tag)" and lists the
unmatched or unfetched rows in one line after the runtime table.
References: nvchecker v2.22 nvchecker_source/github.py (use_max_tag with
include_regex), Renovate's github-tags datasource, GitHub REST "List matching
references"; gh v2.102.0 pkg/cmd/api/pagination.go (paginatedArrayReader).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ng tags from releases
Fixes for the round-3 review findings on 0dadeab7 (no blocking or major defect):
- G1: github_freshness.py records a failed matching-refs call under the record's
matching_tags_errors ({prefix: reason, cut to 160 characters}) and counts it at the
document level as matching_tags_errors, never in partial_errors. A partial error makes
freshness_propose.py blank the repository's drift and trading rows (Inspect AI is a
selected trading card) and holds the propose job. A record carrying it stays pending; the
resume is per repository, so the next run fetches the whole repository again (the
contract's permitted fallback). build_runtime_freshness gives that row
tag_pattern_unfetched with compute_upstream's other fields, and the report blanks nothing.
The fetch's last line also prints the new count.
- G2: a matching tag sets released_at and prerelease to None and drops latest_flag; they
describe the repository's latest GitHub release, which in a monorepo can be another
package's. Dormancy still reads the repository's activity.
- G3: the decision record's first-run table shows Inspect AI compared against the matching
tag 0.3.276 and behind (round-3 local live run on 2026-10-03, 05:29Z to 05:30Z: 488
repositories, 0 errors; 19 entries, 4 behind, 7 not compared) and keeps the 6bc1199 run as
the earlier observation. Verification gives 86 runtime tests, the four registered files of
the round-3 hash failure and the round-3 and round-4 mutation runs; Limits says what a
failed call does and that the list is read with --paginate within the per-call timeout.
- G4: DOTTED_VERSION_RE and the declared pattern compile with re.ASCII, so a tag in
fullwidth digits cannot rank; a missing or uncompilable pattern gives
tag_pattern_unmatched instead of raising; at most 5,000 names are stored per (repository,
prefix), with matching_tags_truncated when a list is cut; the tag-miss sentence and the
stale releases/tags/commit text are corrected.
Tests: 86 runtime tests (6 new, 8 changed). Mutation check: 29 mutants against the 14 new or
changed tests; all 41 mutant/test pairings killed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- docs/decisions/2026-10-02-gpt-runtime-tracking.md: in the observed run, Deep Agents' released_at and prerelease came from its own release of the selected tag, so G2's nulling removed correct values there; it guards the monorepo case. Adds the independent live run at the round-4 code (488 repositories, 0 errors, 4 behind) and the forced tag-list failure check. - README and github_freshness docstring: the resume rule names matching_tags_errors and the per-repository refetch. - freshness_propose: the tag-miss sentence says both reasons keep the release or tag listing as latest; the runtime intro says the last-release and last-commit columns describe the repository's activity, not the tag in the latest column. The report test follows the sentence. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t tag lists and check pointer syntax Fixes for the cross-family review of 2d28e04 (gpt-6-astra at max through native codex exec, read-only sandbox), which returned BLOCK: - X1 (major): github_freshness.py knows, per repository, whether a working file other than runtime-pins.json names it (collect_runtime_only_slugs, compared by normalized slug). The top-level errors and partial_errors, which hold the propose job through freshness_propose.py's upstream-errors.txt and upstream-partial-errors.txt and which saturation_ledger.py reads as the freshness input's completeness, count only those repositories (and any retained record that no working file names), as before runtime-pins.json existed. The failures of the runtime-only repositories (five today) count in the new runtime_only_errors and runtime_only_partial_errors, and runtime_only_repositories lists their slugs. Per-record fields and the resume are unchanged. drift.md's runtime section names, in one line, the rows on a runtime-only repository whose fetch failed; nothing in the drift or propose path reads the new counts. The fetch's last line prints them. - X2 (minor): build_manifest.apply_tag_declaration reads matching_tags_truncated. latest_source becomes tag_pattern_truncated, the selected tag stays the latest (compute_upstream's latest when none matched), and a pinned row is not_compared with reason tag_list_truncated; watch-only and unresolved rows keep their reasons. The tag-miss line names the row with its latest_source, and the "(tag)" marker is not shown for it. The flag is per record, so it applies to every declared prefix of that repository. - X3 (minor): extract_layers.py checks pin_pointer, repository_pointer and a row's array against RFC 6901 (JSON_POINTER_RE) when the declaration is checked, so "tag" for "/tag" raises instead of reading as a moved record. - X4 (nit): the decision record's limits say the runtime rows never trigger a proposal, set drift-status.txt or name receipt component ids, but appear in the committed drift.md; its Decision, Verification and Limits sections cover round 5. The README describes the counters and the cut-list rule. docs/github-automation.md's runtime-table paragraph says that a runtime-only failure does not hold the propose job. Tests: 99 runtime tests, 13 of them new; _fake_gh_api can now fail a repository's primary or releases call. Mutation check: 32 mutants against the 13 new tests; all 57 mutant/test pairings killed. Replay at the reviewed head: a runtime-only 503 wrote 1 to upstream-errors.txt (primary call) or upstream-partial-errors.txt (releases call); with round 5 both say 0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The runtime-only counters cover a repository that only runtime-pins.json names (no foundation, trading, trading-pin or star-candidate working file does), not every repository the runtime table is the only table for. The drift sentence and docs/github-automation.md now say that (round-5 review, minor). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Hot-file protocol (docs/lanes.md): the branch's last commit, on main d2777ee (#619, #620). The PR's own diff, without this registry, is byte-identical to the cross-family-accepted head f6a2f68's diff against its base. scripts/validate.py passed; validate_convergence --all-recorded valid (26 records); the verdict review gate passed against origin/main. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
f6a2f68 to
fb013ea
Compare
…uide is on main The skill-lifecycle bullet pointed to adoption/skills/lifecycle.md "(lands with unit F3)". F3 landed in 3361b34 (#553) and the guide is on main, so the parenthesis is stale. Shared hot file (docs/lanes.md): AGENTS.md and its manifests/evidence.json re-registration are this branch's only commit, rebuilt on main 9b0b8d6 (#634) with AGENTS.md byte-identical to the acknowledged head 1283e4e. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… registry carries main's rows plus the owned blind_checkout.py binding Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…uide is on main (#636) The skill-lifecycle bullet pointed to adoption/skills/lifecycle.md "(lands with unit F3)". F3 landed in 3361b34 (#553) and the guide is on main, so the parenthesis is stale. Shared hot file (docs/lanes.md): AGENTS.md and its manifests/evidence.json re-registration are this branch's only commit, rebuilt on main 9b0b8d6 (#634) with AGENTS.md byte-identical to the acknowledged head 1283e4e. Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Scope
slot, the OpenHands recipepins.json,adoption/sdk/accepted-constraints.txt), and seven watch-only upstreams (pi, oh-my-pi, OpenAI Agents SDK for Python and JS, crawl4ai, Deep Agents, codex-action) are reported against upstream releases with dormancy and archived signals. Upstreams without usable GitHub releases (Inspect AI, codex-action, Deep Agents' monorepo tags) are compared through a declared, anchored tag pattern read with GitHub'sgit/matching-refsand sorted by version; a failed tag list is recorded apart from the repository's partial errors, and a fetch failure on a repository that onlyruntime-pins.jsonnames counts in separateruntime_only_*counters, so neither blanks another row nor holds the propose job. The manifest, the drift table,drift-status.txt, receipts and the propose job are unchanged.e88d59e4(main after Keep daily catalog reports and Monday proposals consistent #639; rebased from56473e4bthroughdcae68bd, the evidence registration redone as the last commit each time)lane:foundation. The branch readsadoption/sdk/accepted-constraints.txtand the install plan and edits neither.tools/sota-convergence/{extract_layers,github_freshness,build_manifest}.py,tools/sota-convergence/README.md,scripts/freshness_propose.py,.github/workflows/catalog-freshness.yml(one flag, one counts line, one upload path),tests/test_catalog_freshness_runtime.py(new),docs/github-automation.md,docs/harness-defaults.md(one anti-pattern row, inserted at the top of the log so it does not meet Keep the root directory as one PATH component; refuse a host HOME of /; six anti-pattern log rows #619's rows at the end),docs/decisions/2026-10-02-gpt-runtime-tracking.md(new).manifests/evidence.jsonre-registrations are the last commit.Relation to open #633 (Codex lane): #633's
catalogs/foundation/runtime-jobs.jsonrecords the source pins its reviews read and has the runtime-worker workflow fetch their metadata; it compares no pin and renders no table. This table compares install and runtime records, so the two are complementary, and once #633 lands its catalog is a pin record onmainfrom whichwatch:pican be promoted.#622 (merged 2026-10-03T05:37Z) carries the
osv-scannerfix for GHSA-vfj7-8cjw-p6xm that had blocked every PR; this branch is rebased onto it.SOTA sources
56473e4b:tools/sota-convergence/extract_layers.pyTRADING_PIN_SOURCESandresolve_trading_pins,build_manifest.pybuild_trading_freshness,scripts/freshness_propose.pyrender_trading_markdown.github_freshness.pymakes; PEP 503 normalized names for the constraints form; RFC 6901 for the pointer form.wf_21fc37c5-123) and not adopted, with the overturn condition in the record: updatecli v0.122.0 (pipeline diff), Renovate 44.132.2 (--platform=local), nvchecker v2.22.Evidence-class table
syntheticpython3 -m unittest tests.test_catalog_freshness_runtime -v: 86 OK; mutation runs (rounds 2-4, 73 mutants) show each new test fails when its behaviour is removeddrift-status.txtunchangedlocal_integrationextract_layers.py,github_freshness.py(488 repositories, 0 errors, 0 partial errors),build_manifest.py --runtime-freshness-out,build_drift_report; table in the decision recordlocal_integrationmatching-refscall; the next run retried and recovered itrepos/openai/openai-python, or 502 on itsreleases/latest) leavesupstream-errors.txt/upstream-partial-errors.txtat 0 and the saturation input completelocal_integrationsource_review(GPT family)gpt-6-astraat max through nativecodex exec, read-only: BLOCK at2d28e04e(four findings, see the PR comment), ACCEPT atf6a2f68eafter round 5local_integrationcmpexit 0mainafter mergeLocal commands run
All at
f6a2f68e(rebased one88d59e4) unless noted; exit codes as returned.Main may move before merge; the branch then rebases, re-registers in its last commit and reruns
validate.pyandvalidate_convergence.py --all-recorded.Decision record
docs/decisions/2026-10-02-gpt-runtime-tracking.md: the gap by row kind, the pin records, ownership (the Codex maintenance task owns upstream detection; this table is a report), the search-first sweep and its overturn condition, the first run, limits and the anti-pattern correction.The sweep's measured gap (Inspect AI not compared because of the name-ordered tag fallback) is closed in this PR. Remaining follow-ups: a registry identity per row through the PyPI simple JSON API (PEP 691, PEP 792 status); advisories through the adopted OSV-Scanner on a generated purl SBOM of the runtime pins; promotion of
watch:piwhen a pi pin record lands onmain(#524 was closed on 2026-10-03; #633 lists pi v1.0.0 as a source pin).Host evidence
Not applicable: no file under
evidence/hosts/changes.Checklist
permissions: contents: read(unchanged).🤖 Generated with Claude Code