Repository navigation
Memory layer S3 r7: symmetric merit rule, maintenance gate, current pins, token ledger, RAG head-to-head (draft, not frozen) - #526
seathatflowsinourveins wants to merge 21 commits into
Conversation
The merit-only head-to-head that decides each host's durable-memory configuration: ai-memory v2.4.0 as the reference arm, 7 challengers, native Claude and Codex memory controls, and frozen reserves. r2 answers all 11 findings of the GPT-6 review of r1: eligibility applies to every configuration, superiority and non-inferiority-plus-cost routes, a qmax-minus-epsilon cost selection, cluster bootstrap statistics, a full dataset and adapter contract, two token ledgers, a defined latency statistic, usefulness lanes, and per-host qualification. Not frozen: it freezes only after a GPT-6 accept on the complete bundle. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
…cement text (not frozen) r2 resolved 9 of r1's 11 findings. r3 applies GPT-6's replacement text for the 2 partly resolved findings and the 4 new ones: - a usefulness floor applied before either replacement route (new high); - a one-sided 5th-percentile lower bound; - one Holm family holding both the superiority and the non-inferiority hypotheses; - a single frozen complete-workload decision ledger, where payload-only counts support only a payload-size claim; - a frozen reference-availability preflight, cluster manifest, seed aggregation, question weighting and RNG. The next GPT-6 review is the bundle review before the freeze. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
…3-prereg-20260927
….4.0 reranker feasibility; v4-based runner rules S3 draft r4 (not frozen): - Reference feasibility from source. Official ai-memory v2.4.0 ships the LLM reranker: AI_MEMORY_RERANKER=llm in crates/ai-memory-cli/src/config.rs:413-420,1744-1759, crates/ai-memory-llm/src/reranker.rs and docs/llm-providers.md:185-192. The docs name no reranker model, so the preregistered qwen3.5-9b-64k stays. Its base is pinned by v4 pins.json:141-147, and the runner verifies the base, the effective template and parameters, and the served blob hashes. - C3' and C4' are defined from the v4 driver's C3/C4. agentmemory's primary arm is its shipped hook path, the A-protocol's D2h shape, with its v4 pins. - Late-candidate rule: later sweep survivors enter only through a dated amendment arm, like reserves admitted after confirmatory execution. The 32-layer sweep is held for the user's Gates A and B. - D-mac-prod: the Mac's production 19b6429 in a C4-shaped configuration, a labelled diagnostic outside the decision family, at the Mac session's request. The 19b6429 / v2.4.0 / v2.4.1 lineage is recorded. - Harness: v4 (#386) is the frozen adapter source, and the S3 runner is new code with the isolation rules from the #386 GPT-6 reviews. The cache proxy's omission is preregistered from A16.3. Only the public reranker Modelfile and its base pin are needed. - A workstation GPU window, a scheduled production stop the user approves; per-host qualification, including reranker timeout latency; Mac Hindsight pending the user's Apple Container decision. - S3's statistics govern, and the A-protocol's do not transfer. Section 12 covers the relationship to A1-A16.3. New blueprints/memory-stack/longmemeval/A17-SUPERSESSION.md (PROPOSED): R1 is not run under A1-A16.3, and S3 supersedes it on both hosts. No confirmatory arm has run on any host, so no look is spent. It takes effect only when the user confirms it at the S3 freeze. No session held an A17 draft (checked with the sessions that had owned it). Every path:line citation in the new text was checked against the files (19 of 19 resolve). Main (#386) is merged in so the v4 citations resolve. validate.py passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
…ls, content-rule clusters, track constructions, agentmemory hook equivalence, fallback rule) The GPT-6 review of r4 (704c9be) returned needs_changes: 2 high, 3 medium, 1 low. r5 applies its replacement text. - high: official v2.4.0 has no embedding-prefix setting (config.rs:239, embedding.rs:337), so C3'/C4' use Qwen3-Embedding-4B's native unprefixed input. They are S3 controls, not reproductions of the prefixed C3/C4, and the bundle verifies the effective embedding-request bytes. - high: clusters follow the canonical (role, content) rule of v4 lme_summarize.py:234. That gives 452 full-track clusters (434 + 18 pairs) and 401 official-track clusters (383 + 18), against 465 and 414 by gold session id. Reproduced here with the frozen v4 clusters() on the pinned dataset (sha256 d6f21ea9..., verified). - medium: the official track is upstream's user-text construction, and the full-session track is the A2 local extension. Both are scored by unchanged evaluate_retrieval. - medium: agentmemory's primary configuration is the explicit local-MiniLM, zero-LLM one (upstream detectEmbeddingProvider returns null by default). The adapter must execute the pinned shipped hooks or prove equivalence, SessionEnd transcripts included. - medium: more than 5% confirmatory C4' fallbacks makes every dependent comparison inconclusive and leaves selection unresolved; there is no post-hoc C3' comparator (A15.2). - low: inputs/ and #379 are described as pending, not present. A17-SUPERSESSION.md's carry-over names the two differences. Section 13 maps the findings. validate.py passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
…puts/ (not frozen) S3 section 2 names the Mac session's 2026-09-27 discovery sweeps as inputs: embedders, rerankers, memory LLMs, generation models and memory systems, each Opus-verified with a Verification section listing every correction and its source. They are committed verbatim from the PR #390 comments, each with a provenance line (comment URL, time, host), under blueprints/memory-layer-s3/inputs/sweep-20260927-mac/. A privacy scan found no personal paths, emails or tokens. The files are registered in manifests/evidence.json. They are discovery inputs only: vendor-reported numbers are motivation for the candidate list, never S3 evidence (section 8). validate.py passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
…rozen) GPT-6 accepted r5 (65fa4fe): all six r4 findings resolved, cluster membership independently matched. Two corrections, neither changing the protocol: - the official track excludes 51 of the 56 single-session-assistant questions and 5 remain (the frozen eligible manifest, checked against the pinned dataset); "the 51" implied all of them; - section 2 records that the Mac discovery inputs are committed (ff12bf1), not pending. validate.py passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
…ss-family review (not frozen) The r6 draft adds the upstream evaluation harnesses (AMB unmodified with frozen MemoryProvider modules, the unchanged LongMemEval eval_utils scorer, scipy and statsmodels statistics, MTEB) and the gateway call envelope for GPT-6 arms. The repair applies all 13 findings of the Claude review (H1-H4, M1-M5, L1-L4): immutable committed input copies under inputs/r6/, K-concurrency with a 96 h coordinator-set window (user may overturn before freeze), the latency preflight, and a new section 15 mapping each finding to its fix. Still DRAFT r6. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Hot-file protocol (docs/lanes.md): main's manifests/evidence.json, the branch's files re-registered, including the five r6 input copies, and both generated reports rewritten with their --write commands. validate.py passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…AMB's retrieval mode; not frozen)
A one-round Claude re-check of the GPT-6 repair found two high defects in the
AMB text the coordinator had directed the repair to adopt, verified here in AMB
03c1d0f1: EvalRunner always constructs a Gemini judge and calls it for AMB's
task_type="open" LongMemEval dataset (runner.py:42, longmemeval.py:67), and
AMB's loader passes gold-revealing "{question_id}_{session_id}" IDs to every
provider (longmemeval.py:328-350). Coordinator decision: one frozen, hashed
LongMemEval dataset adapter passed to unmodified EvalRunner.run in AMB's
retrieval mode (task_type="retrieval": no judge call, runner.py:222-227; opaque
IDs; score_retrieval as the ranked-ID export point), following AMB's own
PrecisionMemBench pattern. Also: cognee v1.6.1 environment versus AMB's 0.5.4
lock, a dispatch wrapper for K concurrency, S3 answerer and judge outside AMB,
validate.py on the pushed tree, attribution and locator precision. Section 16
maps N1-N10; the fixes are not yet independently re-reviewed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Hot-file protocol (docs/lanes.md): main's manifests/evidence.json, the branch's files re-registered, including the five r6 input copies, and both generated reports rewritten with their --write commands. validate.py passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… token ledger, RAG section DRAFT r7 of blueprints/memory-layer-s3/PREREGISTRATION.md, per the coordinator's r7 brief relaying the user's 2026-09-30 requirements. - Maintenance gate (OpenSSF Scorecard Maintained floor): stale = archived or no default-branch commit in 90 days; rerun at freeze and before each run. Stale on 2026-09-30: locomo, LongMemEval (citation only), harbor-datasets, CoIR. - Every system re-pinned at its current release with tag and full commit; ai-memory v2.4.2 is a contender without protected status; reserves engram, Graphiti, supermemory-local; Mem0, Letta, beads out on fit with locators. - Harness: AMB 03c1d0f1 byte-identical; LongMemEval retrieval (retrieval mode) and open-QA (rag mode) adapters, S3 scorer from the paper's metric, thin provider clients in this repo; vendor-bias controls. - Symmetric merit rule replaces the reference and replacement routes: all-pairs Holm (seed 20260927), leader wins, 0.02 band, tie-breaks cost then an exact Scorecard-based repo-quality score. - Token ledger, RAG head-to-head as MTEB 2.21.10 SearchProtocol models, and embedder/reranker selection on disjoint task sets. - r1-r6 review history kept verbatim in Appendix A. Adds the live verification capture (inputs/r7/live-verification-20260930.json) and docs/decisions/2026-09-30-memory-s3-r7.md; re-registers both blueprint files in manifests/evidence.json via host_receipts.register_file. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The GPT-6 cross-family review of DRAFT r7 (cx/gpt-6-astra, effort max) returned needs_changes with seven findings; the coordinator reports 12/12 quoted excerpts matched. This is the single bounded repair round. - F1: static data excludes executed files; harbor-datasets supplies nothing; task-level provenance manifest enforced per trial; S8 regenerates SWE-bench Verified tasks through Harbor v0.23.0's maintained adapter with per-task field identity to SWE-bench_Verified@78f471bf, pending until its amendment. - F2: thin local Hindsight provider keeps the vendor's bank defaults (observations and consolidation on), reads back and matches the effective bank config before ingest, and waits for consolidation; common depth policy (k = 20); parity audit tests async paths and injected failures. Corrects r7's hindsight-all 0.4.17 statement. - F3: maximizer set and for-every-leader quantifier with exact rationals; frozen executable examples, extracted and run, with mutation controls. - F4: C is gross, nonnegative cost of the frozen workload; control difference reported only; equal minima, zero included, go to 7.4; C_RAG. - F5: fallback rate at K and warm latency at concurrency 1 measured separately; the first alone controls stage disabling. - F6: CI term needs a complete, paginated, finished and passing set. - F7: 7.1 item 7, non-inferiority to the no-memory control on S8; the coupling to pending S8 and its alternatives go to the user. Adds inputs/r7/repair1-verification-20260930.json (gh api graphql at pinned commits; REST quota was exhausted) and the verbatim review, and re-registers them with the preregistration in manifests/evidence.json. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fulness and fallback decisions Resolves the independent Opus review of a266a38 (H1 failed-question scoring, H2 usefulness floor, M1-M7, L1-L4) and the GPT-6 re-check's new medium N1, and records two user decisions (2026-09-30) as accepted: - no-winner fallback: the best-scoring eligible system; - usefulness: LongMemEval-V2 (active, @2cc8c54) is the main long-horizon test (beat no-memory, not worse than the best, Holm, multi-seed); MemoryAgentBench TTL/CR is secondary; short code tasks are a do-no-harm check only; MemoryArena (stale since 2026-05-31) is cited only. Also summarizes the eleven candidates' upstream deployment recipes, verified at their pins. Authored on the GPT-6 worker lane (codex exec -p stack-worker, gpt-6.1-sol, effort max); the coordinator reviewed, validated (validate.py passed, evidence manifest check passed) and committed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… open-QA context cap, dependency-scope decision Resolves the Opus cross-family verification of 0d027e0: the no-memory control runs through the same direct evaluation/harness.py entry point with frozen common reader/judge arguments and wire checks; the LME2 resampling unit is fixed shared-trajectory components with a minimum of 20 (else usefulness is pending); open QA receives the shortest prefix covering five sessions, capped at k=20; MemoryAgentBench applicability table; eligible-only best; pending higher scorers keep the fallback unresolved; baseline failures above 5% make usefulness pending; line references corrected. Records user decision 3 (2026-09-30) as accepted: the 90-day rule gates every selected repository; transitive dependencies pass OSV scanning with an SBOM; stale transitive rank_bm25 is flagged without blocking its harness. Authored on the GPT-6 worker lane (gpt-6.1-sol, effort max); coordinator validated and committed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nts and power, HippoRAG 2 to document RAG A 2026-09-30 GPT-6 landscape check recommended seven more memory primaries. The admission rule r7 already has (an official local install and official integrations with both Claude Code and Codex) was applied to each, the same bar the current eight passed: - in, at their latest releases: claude-mem v13.28.0, MCP Memory Service v11.14.0, MemMachine v0.3.9 (spec locators re-read at the release, 31 commits before the extraction head) and ReMe v0.4.1.12; - out on fit: SimpleMem (no Codex path at the release or main); Mem0's 3.3 exclusion stands (Platform routes, no self-hosted OpenMemory in the tree); - routed: HippoRAG 2 to the 8.3 document-RAG head-to-head at a dated default-branch pin, because GitHub v1.0.0 is the legacy HippoRAG 1 code and PyPI's newest (2.0.0a4) predates maintained fixes; RAGFlow as an 8.3 reserve; GitNexus and Serena stay outside S3 (code retrieval is measured in the Harbor token-tool E2E). Memory reserves 4-9 are added in a stated order. All 16 repositories pass the 90-day gate (inputs/r7/landscape-foldin-gate-20260930.json). Counts come from a script that first reproduces the k=8 figures: m 28 -> 66, m_U 36 -> 78, document-RAG Holm pairs 15 -> 21, ingest envelope 3.243B per host, route (a) starts 28,200 per host, and strictest-step power for a 0.05 difference 0.54/0.18/0.08. The text was drafted by the GPT-6 pool lane (cx/gpt-6-astra, max) from a coordinator brief. Coordinator corrections fixed three errors in that brief: the reserve 7/8 order, Serena's route, and a missed eight-system start count. Still draft and not frozen; an Opus check comes next. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…3-r7-20260930 # Conflicts: # manifests/evidence.json
…family check The Claude Opus check of 229307a confirmed: - all 92 cited upstream locators, the counts (k=12, m=66, m_U=78, 21 document-RAG pairs, 3.243B, 28,200 starts and the power figures), the exclusions, the reserve order and the HippoRAG pin exception. It asked for four fixes, applied here by a native gpt-6.1-sol builder at ultra: - MCP Memory Service and MemMachine: the Codex route now cites the actual recipes (multi-client.md:450-477; packages/skills/README.md:13-24 and memmachine-client). The third-party bridge and installer must be pinned and gated before freeze. Codex's native streamable-HTTP client is the labelled alternative. - Serena's Harbor measurement and GitNexus's absence are relayed facts, and the Harbor receipt is required before freeze. S3 has no code-retrieval tool arm; section 8.2's CoIR embedder role is unaffected. - The section 5.5 MemoryAgentBench table gets no-dispatch rows for the four new primaries, and the benchmark's bundled HippoRAG (2.0.0a3) may not stand in for the 398bfdc3 pin. - The section 11 freeze bundle lists both new input captures with their sha256. Also fixed: status lines, supersession pointers, "commit dated" for the lightweight v1.0.0 tag, and the MemMachine re-read tense in the gate receipt. Only screen annotations changed; the captured data is byte-identical. The new receipt sha256 is bed16fa9. Still draft and not frozen. This was the one repair round, with residuals recorded in the PR. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Bounded source reconciliation from the current accepted new-WSL architecture; no benchmark or installation request yet. NativeStack2604 is the user-delegated single final acceptance destination. Current main18eea contains the accepted decision:259–271: durable-memory remains measurement, no protected incumbent, broader S3 candidates/deja-vu at freeze, and LongMemEval-V2 main usefulness against no-memory control. The maintained handbook memory proposal instead carries a four-arm historical proposal with null preregistration_source and LongMemEval-S secondary. Root/Astra source review independently confirms this mismatch; the shorter proposal should not silently become the governing executable comparison. Please reconcile the existing S3 contract's governing scope/candidate coverage and NativeStack2604 target binding with its§11freeze checklist, retaining unresolved inputs/current-pin/host/backend/usage/lifecycle evidence. Your PR526572377d4 remains OPEN/DRAFT/unfrozen and names older workstation/Mac contexts. Root will preserve your owned blueprint paths and review a concrete source update; no fresh memory installation/model benchmark/target config/credential grant is inferred from this handoff. Existing lifecycle-v2 history remains usable within its recorded scope, without implying retrieval-quality or winning-system qualification. Root's compact private reconciliation is |
|
Closed with a record by the PR triage of 2026-10-07 (the command center's ruling, item review-ns2604-coop-20261007T023012Z (the command center's PR-triage ruling of 2026-10-07; proposal by github-ci-finalize, triage-20261007.json)). Not merged; the branch What it holds: blueprints/memory-layer-s3/PREREGISTRATION.md; blueprints/memory-layer-s3/inputs/ (18 files); blueprints/memory-stack/longmemeval/A17-SUPERSESSION.md; docs/decisions/2026-09-30-memory-s3-r7.md Superseded by: Overtaken by the memory-owner decision on main (docs/decisions/2026-10-01-new-wsl-definitive-defaults.md:186, added by ee4ece2 #659: ai-memory 2.5.2, dated 2026-10-03) and by the command center's KEEP ruling (item task-ns2604-coop-20261006T204552Z section 1: no D3r4, LongMemEval or StackMeasure campaign, local head-to-head, new arm or preregistration amendment is revived), recorded in open #800 at af0fbfd:docs/decisions/2026-10-07-memory-keep-and-consolidated-recall.md:46-48,180-181. (confidence: medium: the explicit retirement is in the unlanded #800; main still names PR 526 as the deciding measurement at definitive-defaults.md:342) Reopen trigger: The command center reverses the KEEP ruling, or a source-backed replacement for the memory-owner slot needs a local comparison (definitive-defaults.md:186 names the head-to-head as the overturn path). Reopen with |
Scope
03c1d0f1, with thin providers so that no stale library serves any arm;f7e5e228(Memory layer S3: head-to-head preregistration draft r6 (both hosts; upstream AMB harness in retrieval mode; not frozen) #390 headd161f691+ origin/main merged)lane:foundationblueprints/memory-layer-s3/**,docs/decisions/2026-09-30-memory-s3-r7.md,manifests/evidence.json(hot file, registered throughhost_receipts.register_file)SOTA sources
All pins were verified live on 2026-09-30. Tag-to-commit resolutions, image digests and the gate outputs are in
blueprints/memory-layer-s3/inputs/r7/live-verification-20260930.json.03c1d0f1d27da63034f0931121c858faba512383, used unmodified.mteb/models/models_protocols.py(SearchProtocol).docs/checks.md(the Maintained rule and the check set).2d38dafe5fc4ce20e8716760ba3631f222fd87f0c0bd87c62a4652faa0ca8d1aEvidence-class table
inputs/r7/live-verification-20260930.json(gh api, 2026-09-30)grep -n 9e0b455ffinds citations only (r6 Appendix A is marked superseded)inputs/r7/landscape-foldin-gate-20260930.json(sha256bed16fa9…, gh api + PyPI JSON, 2026-09-30)Local commands run
Decision record
docs/decisions/2026-09-30-memory-s3-r7.md. It records the user's decisions of 2026-09-30 (the ai-memory re-pin, "no bias but evaluate on the sotaness and the quality of repos itself", "NEVER STALED REPOS"), the alternatives and the overturn conditions.Review status (2026-09-30)
229307a1): a GPT-6 landscape check found seven more maintained memory systems.legacy HippoRAG 1 code.
and asked for 4 fixes. The fixes are applied in
4ddea735by a native gpt-6.1-sol builder at ultra.MemMachine (skill, plus memmachine-client in the SBOM).
Open before freeze
blueprints/memory-stack/longmemeval/A17-SUPERSESSION.mdstill describes r6; themanifests/stack.jsonai-memory pin is 2.4.1 (host currency PR).Supersedes the r6 draft in #390.
🤖 Generated with Claude Code