Skip to content

Add a 32-layer landscape-sweep lane to manifest-20260923 and source reviews for its newcomers - #153

Merged
seathatflowsinourveins merged 7 commits into
mainfrom
claude/landscape-sweep-20260923
Sep 24, 2026
Merged

seathatflowsinourveins merged 7 commits into
mainfrom
claude/landscape-sweep-20260923

Conversation

@seathatflowsinourveins

Copy link
Copy Markdown
Owner

Why. manifest-20260923's discovery was thin: 5 web searches across 20 foundation layers, 6 across 12 trading layers and 7 "beyond", with freshness retained (fetched_this_run=0). A current SOTA repository that was never a candidate cannot win the 20260923 re-record, however blind its lanes are. Agreed with agent-lab-17, the tooling owner, with a merge cut-off of 2026-09-24T02:00Z.

Method. Workflow wf_38aa6d5d-d6c, every worker at effort max:

  • Discovery. One Sonnet 5 researcher per layer for all 32 layers. Budget per layer: at most 10 WebSearch, 4 WebFetch and 30 gh api calls. Actual totals: 160 web searches, 61 fetches and 871 gh calls.
    • The researcher proposes at most 4 repositories that are not among the 259 catalog-known ones (forks, mirrors and renames resolved), that serve the layer's requirement and are maintained.
  • Refuters. Two per layer, both defaulting to refuted: a facts/identity lens (Sonnet 5) and a fit/standing lens (Opus 5.5). A candidate survives only when neither refutes it.
  • Labels. New names are only keep_but_compare or targeted_candidate, never promoted (the generator's PROPOSABLE_LABELS).

Result. 98 proposals: 11 survivors and 87 refuted, all published with their votes. The survivors:

  • document-retrieval: datalab-to/marker, opendatalab/MinerU
  • web-research: exa-labs/exa-mcp-server
  • ci-supply-chain: ossf/scorecard, renovatebot/renovate
  • observation-inference: grafana/tempo
  • secrets-credentials: trufflesecurity/trufflehog
  • git-github-automation: github/gh-aw
  • research-factors-ml: lightgbm-org/LightGBM, microsoft/RD-Agent
  • security-supply-chain: google/osv-scanner

Build. build_manifest.py is unchanged. I reproduced the committed manifest byte for byte from the private work dir sota-convergence-20260923 (lanes-with-critic.json, the committed citation review, both original --checkout-root values). I then appended this lane (new files beside the originals; the originals are not modified) and rebuilt.

  • 0 existing candidates, components or entries change. The manifest changes only in: +98 candidates, lane_calls/lane_limits for this lane, and counts.candidates_total/candidates_by_disposition (73 → 171).

Evidence for the blind lanes. Each survivor has a neutral evidence/artifacts/landscape-sweep-20260923/<owner>-<repo>.json (evidence_class: source_review) containing:

  • the pinned upstream commit, SPDX license and README path;
  • the repository description and verbatim README excerpts at that commit;
  • a reviewer summary.

Popularity, recency, catalog-status and comparison sentences are filtered out, and the layer's current winners are never named. Each file is registered in manifests/evidence.json files[] and listed first in its candidate's evidence[], so agent-lab-17's --manifest-newcomers packets attach it.

Checks (local, raw).

  • validate.py, evidence_manifest --check, landscape.py and build_verdicts --check pass.
  • new_host_grand_list --check and component_matrix --check pass.
  • verdict_review_gate --base origin/main passed.
  • 567 tests OK: sota_convergence, layer_verdicts, catalog_freshness_propose, verdict_review_gate, lane_packets and landscape.
  • Guarded gitleaks: no leaks.

Limits.

  • Source review and metadata only. No candidate was installed, run or compared with a winner. Survival means the proposal withstood fact and fit checks, not superiority.
  • upstream_now is as observed during the run.
  • Call counts cover the discovery agents only.
  • A first run at lower worker effort was stopped at 7 of 32 returns when the effort policy changed, and its results were discarded.
  • The completeness critic's follow-up round is still running. If it lands before the cut-off it will be a second commit; otherwise it becomes a dated input for the next wave.

🤖 Generated with Claude Code

…osals, 11 survivors with source reviews

A per-layer sweep of all 32 layers (one researcher per layer, a facts/identity and a fit/standing
refuter per layer, every worker at effort max) proposed 98 repositories the catalog did not know;
11 withstood both refuters. The manifest is rebuilt with the unchanged generator from the original
private inputs plus this lane (the original inputs alone reproduce the committed manifest byte for
byte). Survivors carry a neutral source_review file at a pinned upstream commit, registered and
listed first in their evidence, for the blind re-record lanes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-24T00:31:04.146515Z bf08e09 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: bf08e09581

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread catalogs/sota-convergence/manifest-20260923.json Outdated
Comment thread catalogs/sota-convergence/manifest-20260923.json Outdated
Comment thread catalogs/sota-convergence/manifest-20260923.json Outdated
Comment thread catalogs/sota-convergence/manifest-20260923.json Outdated
Comment thread evidence/artifacts/landscape-sweep-20260923/datalab-to-marker.json
Comment thread catalogs/sota-convergence/manifest-20260923.json Outdated
…ceipts, existing newcomers covered

- F2: repositories already recorded in catalogs/foundation/automation.json (active osv-scanner and
  Scorecard lanes; dated considered_not_activated decisions for Renovate and gh-aw; harden-runner)
  are known, not candidates; the two active ones become open gaps on their layers instead.
- F1/F3/F4: receipts carry only the upstream's own words (repository description and verbatim
  README excerpts at the pinned commit) and only the pinned tree/README URLs; popularity, standing,
  recency, status and comparison sentences are filtered; winner names match whole words; nav-only
  excerpts and user/sponsor sections are skipped.
- agent-lab-17's request: the manifest's 57 existing surviving newcomers (51 repositories, 6 of them
  Hugging Face models) get the same neutral source-review files, prepended to their evidence.
- gitleaks: a rule-scoped, exact-file, whole-line allowlist for the manifest's "pin" git commit ids,
  which the sourcegraph-access-token rule's keyword activated once lane text named Sourcegraph; with
  regression tests that other fields and other files stay detected.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins seathatflowsinourveins changed the title Add a 32-layer landscape-sweep lane to manifest-20260923 (98 proposals, 11 survivors) for the re-record Add a 32-layer landscape-sweep lane to manifest-20260923 and source reviews for its newcomers Sep 24, 2026
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Update at 703533c9, after the independent review (fix_required: two mediums and two lows, all addressed).

  • F2, known repositories. The sweep's known set lacked catalogs/foundation/automation.json.
    • Scorecard, Renovate, harden-runner, gh-aw and osv-scanner are recorded there, either as active interfaces/lanes or as dated considered_not_activated decisions. They are now known, not candidates.
    • The two active ones (osv-scanner PR check, Scorecard lane) are recorded as open gaps on security-supply-chain and ci-supply-chain: the layer rows omit automation this repository already runs.
    • Result: 93 proposals, 7 survivors (marker, MinerU, exa-mcp-server, tempo, trufflehog, LightGBM, RD-Agent), 86 refuted.
  • F1/F3/F4, receipts.
    • Receipts now carry only the upstream's own words: the repository description and verbatim README excerpts at the pinned commit.
    • Sources are only the pinned tree and README URLs.
    • The filter is wider (standing, comparison and count claims), winner names are matched as whole words, and navigation-only excerpts and user/sponsor sections are skipped.
  • Added at agent-lab-17's request. The manifest's 57 existing surviving newcomers (51 repositories, 6 of them Hugging Face models) get the same neutral source-review files under evidence/artifacts/source-review-20260923/. Each path is prepended to the entry's evidence[] in a copy of the lane record. Those 57 entries change by exactly that one path.
  • gitleaks.
    • CI's gitleaks dir flagged 5 sourcegraph-access-token false positives: the manifest's existing "pin" commit ids, activated once lane text named Sourcegraph (the SCIP format's former sourcegraph/scip home).
    • A new allowlist entry is scoped to that rule, the exact file, and a whole anchored "pin" line holding a 40-hex id with an optional <version> @ prefix.
    • Regression tests test_d/test_d2 check that other fields in the file and pin lines in other files stay detected. Both were mutation-checked.

Local checks.

  • validate.py, evidence_manifest --check, landscape.py, build_verdicts --check, the grand-list and matrix --check, and the gate all pass.
  • 624 tests OK.
  • Full gitleaks dir: no leaks. gitleaks git over the range: no leaks.

A re-review is running.

…is a mirror; backtrader unpushed since 2024-08

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…er, no duplicate backtrader gap

- L1: a sentence dropped between two kept ones is marked [...] so a joined excerpt never reads as
  contiguous upstream text.
- L2: counts and superlatives outside the word-boundary filter ("1B+", "#1", "most complete",
  "leaderboard") are dropped.
- L3: the backtrader gap is removed; its home layer already carries the beyond lane's
  unmaintained_signal.
Pinned commits are unchanged (reused from the reviewed receipts); 10 receipts change in text only.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Codex P1 (RD-Agent): its comparison scores on a final segment the feedback loop never sees.
- Codex P2 (TruffleHog, 2 layers): the verification arm uses a controlled, revocable credential.
- Codex P2 (MinerU): stale v1.0-era OmniDocBench figures and CJK framing removed.
- Codex P2 (Marker): model-weight license terms recorded separately from the code license.
- Codex P1 (stopped run): the first run's 9 discovery returns are retained, label-free, with call
  counts and whether the completed run re-proposed each repository; its usage is unknown.
- The completeness critic's follow-up round (same two-refuter rule) is merged: 107 proposals,
  11 survivors (adds PaddleOCR, ollama, betterleaks, claude-code-action), with prior documentary
  records disclosed; completed-run usage recorded (119 agents, 12.1M subagent tokens).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tion records disclosed

- M1: PaddleOCR-VL-1.6 is 3rd on OmniDocBench v1.6 (README at f133a71e9e), not 1st; its proposed
  label is lowered to keep_but_compare because its fit vote called the rank-1 basis void.
- M2: claude-code-action's prior records are cited correctly (targeted_candidate with an executed
  CI smoke arm in the 2026-09-22 SDK sweep; decision HOST-09) in the lane limits and its comparison.
- L1: RD-Agent keeps its reproduce-the-published-advantage stay condition.
- L2/L3: a stale Release claim is corrected and a private scratch name is removed from evidence.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… decision (HOST-09 is keep-but-compare)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins merged commit 9eac1f9 into main Sep 24, 2026
22 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the claude/landscape-sweep-20260923 branch September 24, 2026 01:18
seathatflowsinourveins added a commit that referenced this pull request Sep 24, 2026
…h merged evidence (#170)

Records only, no gate status change: runtime-target IBKR cites the passed 1.231.0 paper receipt (#147) while rc5 local acceptance stays not_established (blocker nautilus#4983); dashboard checkpoint cites the 32-layer sweep (#153) and the #162 mover research result. Independently reviewed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant