Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions catalogs/sota-convergence/manifest-20260923.json
Original file line number Diff line number Diff line change
Expand Up @@ -17460,10 +17460,10 @@
"Discovery: one Sonnet 5 researcher per layer at effort max (WebSearch <=10, WebFetch <=4, gh api <=30 per layer); adversarial verification: a facts/identity refuter (Sonnet 5, max) and a fit/standing refuter (Opus 5.5, max) per layer, each defaulting to refuted; a candidate survives only when neither refutes it; one Opus 5.5 (max) completeness critic and one bounded follow-up round over at most 8 critic-flagged layers.",
"Source review and GitHub metadata only: no candidate was installed, run, benchmarked or compared with a winner; survival means the proposal withstood fact and fit checks, not that the candidate is better than a winner.",
"upstream_now values are as observed by the researcher's gh api calls during the run (2026-09-23/24 UTC), not a github_freshness.py fetch.",
"Call counts cover the discovery agents only; refuter and critic calls are not counted here. Completed run wf_38aa6d5d-d6c: 119 agents (32 discovery, 31+31 refuters, 1 critic, 8 follow-up discovery with 8+8 refuters), 12,123,864 subagent tokens, 2,823 tool uses, 6,615 s wall clock, as reported by the workflow runtime.",
"Call counts cover the discovery agents only; refuter and critic calls are not counted here. Completed run wf_38aa6d5d-d6c: 119 agents (32 discovery, 31+31 refuters, 1 critic, 8 follow-up discovery with 8+8 refuters), 12,123,864 subagent tokens, 2,823 tool uses, 6,615 s wall clock, as reported by the workflow runtime; per-request usage from child-usage.mjs (a different counter, not to be summed with it): all 119 children resolved to claude-sonnet-5 or claude-opus-5-5 at effort max, 4,225,386 output / 3,544 input / 138,862,781 cache-read / 9,733,523 cache-creation tokens (retained at evidence/artifacts/landscape-sweep-20260923-attempts/child-usage-wf_38aa6d5d-d6c.json).",
"The completeness critic's follow-up round (8 critic-flagged layers, same two-refuter rule) is merged into this lane with the first round.",
"Prior records of survivors (no adopt or reject decision, so they remain candidates): PaddlePaddle/PaddleOCR is discovery_only in catalogs/landscape/claude-independent-discovery.json, where a 2026-09-21 refuter recorded PP-StructureV3 below MinerU2.5 on OmniDocBench; anthropics/claude-code-action is targeted_candidate (3 votes) in catalogs/sota-convergence/sdk-runtime-coverage-20260922.json with an executed CI smoke arm (/gated_items/rows/0), and decision HOST-09 (docs/harness-rules-convergence-20260922.md) keeps it unadopted until a pinned-commit check of the action repository is run.",
"A first run (workflow wf_28d47bbc-1b5) at lower worker effort (discovery Sonnet high, refuters Sonnet medium and Opus high) was stopped after 9 of 32 discovery returns and no refuter returns, when the effort policy changed to max for every worker. It is retained, not merged, at evidence/artifacts/landscape-sweep-20260923-attempts/wf_28d47bbc-1b5.json with each returned repository, its call counts, and whether this run re-proposed it; its provider usage was not captured (unknown).",
"A first run (workflow wf_28d47bbc-1b5) at lower worker effort (discovery Sonnet high, refuters Sonnet medium and Opus high) was stopped after 9 of 32 discovery returns and no refuter returns, when the effort policy changed to max for every worker. It is retained, not merged, at evidence/artifacts/landscape-sweep-20260923-attempts/wf_28d47bbc-1b5.json with each returned repository, its call counts, and whether this run re-proposed it. Its usage, measured afterwards from the retained transcripts by agent-lab's .claude/workflows/child-usage.mjs: 17 discovery children (9 returned, 8 stopped in flight), claude-sonnet-5 at effort high, at least 168,052 output / 302 input / 8,233,147 cache-read / 773,809 cache-creation tokens. The tool reports the run incomplete, so these are lower bounds and the unobserved remainder is unknown; its output is retained at evidence/artifacts/landscape-sweep-20260923-attempts/child-usage-wf_28d47bbc-1b5.json.",
"Post-review amendments (Codex cross-family review of #153, 2026-09-24), votes and survival unchanged: RD-Agent's comparison now scores on a final segment its feedback loop never sees; both TruffleHog comparisons use a controlled, revocable credential for the verification arm; MinerU's stale v1.0-era OmniDocBench figures and CJK framing are removed from its gap and comparison. Final review of #153: PaddleOCR's rank-1 OmniDocBench claim is corrected to 3rd (README at f133a71e9e) and its proposed label lowered from targeted_candidate to keep_but_compare, because its fit vote's text called the rank-1 basis void while recording refuted=false; claude-code-action's comparison cites its prior records and its Release evidence is corrected; RD-Agent keeps both stay conditions."
]
},
Expand Down
Loading
Loading