Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 8 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,10 +4,17 @@ A current, evidence-backed reference for native **Codex + Claude Code**, with sc

**Snapshot: September 20, 2026.** This repository records a tested selection and its limits. It is not a claim that every available framework is installed or that an independent universal SOTA benchmark has been won.

The new **[evidence-led convergence practice](blueprints/convergence-practice/README.md)**
connects public own repositories, stars and curated-list discovery to pinned
source review, frozen experiments and explicit adoption decisions. It adds an
executed external retrieval comparison and a reusable acceptance protocol for
native agent work, recovery and local inference. Research findings remain
separate from a new host's runtime acceptance.

The **[US-equities grand catalog](catalogs/us-equities/README.md)** now covers
**147 unique repositories in 152 layer decision cards**, **20 model entries**,
and an auditable **342-star coverage ledger**. Its combined index includes
**502 repository identities** across all 342 public stars and 160 beyond them,
**504 repository identities** across all 342 public stars and 162 beyond them,
with typed, validated pointers to decisions and evidence. The latest
[architecture research wave](catalogs/us-equities/architecture/README.md) adds
source reviews, awesome-list coverage and official Alpaca constraints.
Expand Down
153 changes: 153 additions & 0 deletions blueprints/convergence-practice/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,153 @@
# Evidence-led convergence practice

This practice turns repository discovery into a tested change to a native agent
workflow. Its target is reliable issue-to-result work across a coordinator and
explicit worker hosts, with source-grounded research, bounded context, native
authentication and recoverable execution.

The September 20, 2026 wave builds on the existing
[native stack](../../README.md), [source-review union](../../catalogs/us-equities/decision-index.json),
[research runtime](../us-equities/research-runtime/README.md) and
[measured retrieval baseline](../us-equities/retrieval-evaluation/README.md).
It does not replace those receipts or transfer their acceptance to a new host.
SOTA is the research direction; this publication makes no universal best-stack
or end-to-end autonomous-system claim.

## This wave's evidence

The [source review](../../catalogs/convergence-practice/source-review.md) refreshed
one public owned repository and all **342 public stars**. All stars were already
represented in the prior public index. Three pinned awesome lists supplied
selected discovery links; seven candidates received pinned README, license and
supporting-source review. The ledger retains 26 selected source-file hashes.
Harbor is a new candidate; llama.cpp was already a maintained trial and is newly
represented in the public index. Neither becomes an installed default here.

The [machine-readable source ledger](../../catalogs/convergence-practice/source-review.json)
distinguishes the source-review wave from subsequent execution. Its public
inventory is allowlisted metadata; none of the private ecosystem's source,
histories, accounts or machine configuration is included.

The [external retrieval experiment](arb-trace2code/README.md) runs two unmodified
upstream baselines against a complete published failure-trace subset. Its inputs,
exact implementation pin, observed results and replay instructions are retained.
It evaluates file retrieval, not native agent repair, answer correctness or
provider savings.

The [local acceptance fixture](local-fixture/fixture.json) freezes 24 questions
and 15 public documents at an immutable repository revision: 20 answerable
questions plus four corpus-specific no-answer cases. Its
[freeze receipt](local-fixture/freeze-receipt.json) predates ranking. The author
read the sources, so this is a transparent integration fixture, not blinded
external relevance judgment. Preserve these exact bytes for any comparison.
The [executed local replay](local-fixture/evaluation.md) retains all 48 rankings,
the denominator rules and a fixture-specific recipe using unchanged upstream APIs.

## What the comparisons changed

The external failure-trace subset favors ARB's path/symbol-aware lexical
implementation on mean Recall@20: **69.64% versus 49.34%** for its BM25 baseline.
The separate local documentation fixture favors BM25 on exact first-source
retrieval: **95% versus 80%** across the 20 answerable questions. Both methods
reach 100% Recall@3 on that small local fixture. These are different corpora,
labels and cutoffs; their scores must not be pooled into one leaderboard.

Both local rankers also return documents for all four no-answer questions.
That observation is not an answer-generation failure rate: no answer model or
abstention policy was evaluated. It demonstrates why a ranking alone cannot
certify that the requested information exists in the corpus.

The resulting decision is to retain the native production choices and adopt
the replayable evaluation practice. There is no evidence here for replacing
QMD, SocratiCode, memory or a native agent harness. A proposed replacement must
first beat the relevant current workflow on its intended task and pass its
integration and recovery checks.

## The reusable loop

1. **Name the failure and the useful outcome.** Start from a real task, a measured
miss or an explicit missing capability. Define success before selecting a tool.
2. **Discover broadly, review narrowly.** Refresh public own repositories and
stars, follow relevant curated-list links, then inspect primary upstream
documentation and source. Record how each candidate was discovered. An awesome
list is a discovery source, not proof of compatibility or quality.
3. **Challenge the current choice.** Compare the existing native workflow,
a minimal alternative and the proposed addition. State what each overlaps,
its license at the exact pin, hardware/dependency needs and expected benefit.
4. **Freeze the experiment.** Pin source and input hashes, questions, relevance
judgments, budgets and evaluation rules before running. A question authored
from source is a repository-specific fixture, not an independent user study.
5. **Run in isolation.** Use an explicit project and disposable environment.
Retain successful results, failures and skipped work. Keep credentials in
native stores and raw machine details out of the publication.
6. **Decide from the result.** Adopt only the capability and host scope actually
accepted. Retain the current implementation when the candidate has no measured
advantage. An inconclusive result is useful evidence and leaves the gate open.
7. **Publish and replay.** Include commands, input/code hashes, metric definitions,
limitations, rollback and the next unresolved test. Check the exact publication
commit in CI. Revisit after a relevant upstream, model, corpus or host change.

[protocol.json](protocol.json) makes the lanes and required evidence explicit.
The contract adds no scheduler, always-loaded instructions, proxy, memory store or
runtime service. Model calls, package installation and broker activity remain
separate actions with their own scope.

## Keep three decisions separate

| Question | Evidence that can answer it |
| --- | --- |
| Is the project worth investigating? | Public discovery provenance and primary-source review |
| Does this pinned implementation work here? | A useful execution on the named platform, with inputs and observed output |
| Does it improve our work? | A matched task comparison with correctness, resource use and failure behavior |

Discovery, source review, offline artifact checks, native CLI behavior,
model-mediated task outcomes and recovery acceptance are distinct evidence
classes. They are not interchangeable badges. Record an adoption decision
separately: **retain**, **trial**, **adopt within scope**, **defer**, or **reject**.
State which fact would change that decision.

## Next acceptance lanes

| Lane | Small useful acceptance | What remains unproved by a version check |
| --- | --- | --- |
| Retrieval | Frozen source questions, exact identifiers, negative cases, and a larger external benchmark | Answer correctness, safe abstention and successful code edits |
| Native agent engineering | One real issue, isolated change, relevant tests and independently reviewed patch | General autonomous repair or production safety |
| Memory | Retrieve an explicitly approved decision in its project and refuse a different scope | Automatic consolidation quality or safe transcript capture |
| Access and recovery | Reconnect after interruption, preserve job identity, cancel descendants and recover selected state | Cold boot, network changes and restoration on an independent host |
| Local inference | One defined workload compared with the current baseline; record quality, memory and latency | Superiority from model size, release date or tokens per second alone |
| Research application | Reproduce a source-to-deterministic-result packet with time/provenance checks | Historical data entitlement, strategy validity or order authority |

The existing financial research path keeps model reasoning separate from numeric,
risk and order state. Its [research protocol](../us-equities/acceptance-wave/research-protocol.md)
and [open gates](../../catalogs/us-equities/convergence-review.json) remain
authoritative for that application; this practice does not advance those gates.

## Measure the whole task deliberately

Report retrieval coverage separately from answer or patch correctness. Specify
exact-match versus partial-path scoring, top-k denominators, ties, empty results
and errors. Never silently score a failed search as a correct abstention.

Record warm/cold conditions, dependency versions and whether timing includes
startup. Selected text bytes, tokenized context, cache hits, provider usage and
billing are different quantities. Claim net token or cost savings only from a
matched task comparison with complete usage categories and comparable quality.

For new candidate models or retrieval methods, keep a frozen evaluation set
separate from development examples. Once results have influenced changes, that
set is a regression fixture; use fresh questions for the next generalization
claim. Avoid turning repeated review by similar agents into an independent
benchmark or a statistical confidence claim.

## Public boundary and continuation

Publish only public source identities, content hashes, sanitized measurements,
reviewed code and reproducible instructions. Private repositories can consume
this practice locally without exporting their source, prompts, accounts, memory
or operational topology. A public star list is a dated discovery snapshot, not
blanket approval to install its contents.

Continue by selecting one open lane and the smallest experiment that could change
its decision. Keep the accepted implementation and rollback available while a
candidate is tested. A new release triggers inspection; it does not automatically
change a working pin.
132 changes: 132 additions & 0 deletions blueprints/convergence-practice/arb-trace2code/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,132 @@
# ARB trace2code: exact-release replay

This is an actual offline replay of the complete **101-case `v2_trace2code`**
release using Agent Retrieval Bench's unmodified lexical and BM25 evaluators.
It is one positive-only failure-trace task, not an evaluation of installed
QMD, rg, SocratiCode, an embedding model, or an agent's repair success.

| Native ARB ranker | Samples | Skipped | Recall@5 | Recall@10 | Recall@20 | MRR |
|---|---:|---:|---:|---:|---:|---:|
| Lexical | 101 | 0 | 0.343234 | 0.481848 | 0.696370 | 0.207453 |
| BM25 | 101 | 0 | 0.222772 | 0.321782 | 0.493399 | 0.163848 |

Recall is the arithmetic mean of each sample's fraction of gold file paths
retrieved by the cutoff. It is **not** the fraction of solved cases or hit@k.
MRR uses the first exact gold-file rank across the **full file ranking**, not
only its first twenty files. There are 80 cases with one gold file, 16 with two,
and five with three, for 127 gold-file references. File ranks deduplicate native
chunk rankings by first occurrence of each case-sensitive repository path.

The seven repository counts are caddy 7, clap 1, etcd 4, gin 56, click 26,
pytest 1 and tokio 6. Gin and Click together contribute 82/101 cases. No
repository balancing, confidence interval, significance test, parameter tuning,
or claim of general superiority is made. Per-repository raw means and paired
differences are retained in [metrics.json](metrics.json).

## Pins, algorithms and licensing

- Evaluator: [`v0.2.1`, commit `b487f3866cc13dd971819cb902517a6a50282404`](https://github.com/eyuansu62/agent-retrieval-bench/tree/b487f3866cc13dd971819cb902517a6a50282404).
- Dataset: [revision `5901e1ee3aff048290db72edf9c63bc498b79ea3`](https://huggingface.co/datasets/eyuansu71/agent_retrieval_bench/tree/5901e1ee3aff048290db72edf9c63bc498b79ea3), release `v2_trace2code`.
- Archive SHA-256: `19b252e8cfff42107fedc74005dbb6972f2970af33651ce0c1571546819e41c4`.
- The archive contains 98 frozen base-commit corpora, 24,883 snapshot file
records and 356,074 snapshot chunk records. Repeated files across snapshots
are counted repeatedly. Its compressed size is 39,295,446 bytes and its
extracted file bytes total 300,175,960.
- The evaluator has no mandatory third-party runtime dependency. This run used
a separate stdlib-only Python 3.14.7 environment on macOS arm64.

ARB lexical scoring combines token document frequency and query frequency with
full-path (+25), basename (+8) and symbol (+5) substring bonuses, then divides
by the square root of the number of unique chunk tokens. BM25 retains the
upstream `k1=1.5`, `b=0.75` and query-frequency weighting. Both tokenize camel-case
split, lowercase ASCII alphanumeric text containing path, symbol, kind and
content. Ties use path then chunk identifier. Neither implementation is an
rg or QMD baseline. [Pinned scoring source](https://github.com/eyuansu62/agent-retrieval-bench/blob/b487f3866cc13dd971819cb902517a6a50282404/src/agent_retrieval_bench/baseline.py).

The evaluator, benchmark metadata and reports are MIT licensed. Corpus content
retains the licenses of its original repositories; no corpus or query text is
redistributed here. See [upstream data licensing](https://github.com/eyuansu62/agent-retrieval-bench/blob/b487f3866cc13dd971819cb902517a6a50282404/DATA_LICENSE.md).

The first run used the catalog's retained source snapshot `07014c986f3deadb1548c62b32c0ffbe6a81465d`,
which declares version 0.2.1 but is not the release-tag commit. After verifying
the tag discrepancy, both rankers were rerun at the exact release commit.
The complete evaluator source tree is byte-identical at the two commits
(`0fd0466fb3e24c40c0bdf9468872bea9c9a61fa1`); the differences are documentation
and citation changes. Both executions have identical per-sample detail bytes
and metrics. The initial observation is preserved in the receipt's provenance,
not silently relabeled as an exact-tag run.

## Retained evidence and limits

[receipt.json](receipt.json) records source/data pins, parameters, environment,
exit codes, observations, checks and artifact hashes.
[input-files.json](input-files.json) hashes every extracted input file;
[source-files.json](source-files.json) hashes evaluator code and license sources.
[paired-samples.json](paired-samples.json) retains only public sample identifiers,
repository/base commits, gold paths and each ranker's exact gold ranks, allowing
independent Recall/MRR recomputation without republishing queries or code chunks.

The two `upstream-*-summary.json` files preserve native output unchanged.
**Their legacy `@8k` fields are not tokenizer-measured BCY or token savings.**
The pinned baseline uses Python `len(text)` characters, and accepts an oversized
first chunk, so it does not even enforce a strict 8,000-character cap in that
case. Those legacy fields are excluded from headline metrics. No abstention,
answer correctness, span quality, model inference, provider usage or billing
claim follows from this positive-only file-retrieval experiment.

The exact-release runs were sequential, single warm observations after the
earlier source-identical run. Timing is informational; it is not a controlled
latency benchmark. Both exact-release commands exit 0; schema validation finds
101 valid samples and zero invalid samples. Fifteen upstream corpus/baseline
tests pass. No runtime/default adoption is implied.

## Replay the upstream evaluator

Use a new empty directory and an existing Python 3.14.7 executable as
`ARB_PYTHON`. This recipe creates only that directory's environment/data; it
does not install a global tool, register a service or contact a model provider.

```sh
git clone https://github.com/eyuansu62/agent-retrieval-bench.git upstream
git -C upstream checkout --detach b487f3866cc13dd971819cb902517a6a50282404
"$ARB_PYTHON" -m venv --without-pip runtime
mkdir data results
curl --fail --location --max-time 120 \
--output agent_retrieval_bench_v2_trace2code.tar.zst \
https://huggingface.co/datasets/eyuansu71/agent_retrieval_bench/resolve/5901e1ee3aff048290db72edf9c63bc498b79ea3/releases/v2_trace2code/agent_retrieval_bench_v2_trace2code.tar.zst
runtime/bin/python - <<'PY'
import hashlib
import tarfile
from pathlib import Path, PurePosixPath

path = Path("agent_retrieval_bench_v2_trace2code.tar.zst")
with path.open("rb") as handle:
actual = hashlib.file_digest(handle, "sha256").hexdigest()
if actual != "19b252e8cfff42107fedc74005dbb6972f2970af33651ce0c1571546819e41c4":
raise ValueError("Archive checksum mismatch")
with tarfile.open(path, "r:zst") as archive:
members = archive.getmembers()
if len(members) != 120 or sum(item.size for item in members) != 300175960:
raise ValueError("Archive size/member inventory mismatch")
for item in members:
name = PurePosixPath(item.name)
if name.is_absolute() or ".." in name.parts or not (item.isfile() or item.isdir()):
raise ValueError("Unsupported archive member")
archive.extractall("data", members=members, filter="data")
PY
PYTHONPATH=upstream/src runtime/bin/python -m agent_retrieval_bench.cli \
validate data/benchmark/v2_trace2code/samples.jsonl
for ranker in lexical bm25; do
PYTHONPATH=upstream/src runtime/bin/python -m agent_retrieval_bench.cli \
eval-baseline data/benchmark/v2_trace2code/samples.jsonl \
--corpus data/corpus/v2_trace2code --ranker "$ranker" \
--candidate-filter all_files --no-keep-list --no-progress \
--out "results/$ranker-summary.json" \
--details "results/$ranker-details.jsonl"
done
```

Run from the directory containing `data`: the released manifest's chunk paths
are relative to that directory. Compare metrics and detail hashes, not elapsed
time fields. Do not substitute `--dry-run`, an answer-only candidate list, or a
sample limit for this complete released-subset replay.
Loading