Repository navigation
Memory layer S3: head-to-head preregistration draft r6 (both hosts; upstream AMB harness in retrieval mode; not frozen) - #390
seathatflowsinourveins wants to merge 11 commits into
Conversation
The merit-only head-to-head that decides each host's durable-memory configuration: ai-memory v2.4.0 as the reference arm, 7 challengers, native Claude and Codex memory controls, and frozen reserves. r2 answers all 11 findings of the GPT-6 review of r1: eligibility applies to every configuration, superiority and non-inferiority-plus-cost routes, a qmax-minus-epsilon cost selection, cluster bootstrap statistics, a full dataset and adapter contract, two token ledgers, a defined latency statistic, usefulness lanes, and per-host qualification. Not frozen: it freezes only after a GPT-6 accept on the complete bundle. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
…cement text (not frozen) r2 resolved 9 of r1's 11 findings. r3 applies GPT-6's replacement text for the 2 partly resolved findings and the 4 new ones: - a usefulness floor applied before either replacement route (new high); - a one-sided 5th-percentile lower bound; - one Holm family holding both the superiority and the non-inferiority hypotheses; - a single frozen complete-workload decision ledger, where payload-only counts support only a payload-size claim; - a frozen reference-availability preflight, cluster manifest, seed aggregation, question weighting and RNG. The next GPT-6 review is the bundle review before the freeze. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
…3-prereg-20260927
….4.0 reranker feasibility; v4-based runner rules S3 draft r4 (not frozen): - Reference feasibility from source. Official ai-memory v2.4.0 ships the LLM reranker: AI_MEMORY_RERANKER=llm in crates/ai-memory-cli/src/config.rs:413-420,1744-1759, crates/ai-memory-llm/src/reranker.rs and docs/llm-providers.md:185-192. The docs name no reranker model, so the preregistered qwen3.5-9b-64k stays. Its base is pinned by v4 pins.json:141-147, and the runner verifies the base, the effective template and parameters, and the served blob hashes. - C3' and C4' are defined from the v4 driver's C3/C4. agentmemory's primary arm is its shipped hook path, the A-protocol's D2h shape, with its v4 pins. - Late-candidate rule: later sweep survivors enter only through a dated amendment arm, like reserves admitted after confirmatory execution. The 32-layer sweep is held for the user's Gates A and B. - D-mac-prod: the Mac's production 19b6429 in a C4-shaped configuration, a labelled diagnostic outside the decision family, at the Mac session's request. The 19b6429 / v2.4.0 / v2.4.1 lineage is recorded. - Harness: v4 (#386) is the frozen adapter source, and the S3 runner is new code with the isolation rules from the #386 GPT-6 reviews. The cache proxy's omission is preregistered from A16.3. Only the public reranker Modelfile and its base pin are needed. - A workstation GPU window, a scheduled production stop the user approves; per-host qualification, including reranker timeout latency; Mac Hindsight pending the user's Apple Container decision. - S3's statistics govern, and the A-protocol's do not transfer. Section 12 covers the relationship to A1-A16.3. New blueprints/memory-stack/longmemeval/A17-SUPERSESSION.md (PROPOSED): R1 is not run under A1-A16.3, and S3 supersedes it on both hosts. No confirmatory arm has run on any host, so no look is spent. It takes effect only when the user confirms it at the S3 freeze. No session held an A17 draft (checked with the sessions that had owned it). Every path:line citation in the new text was checked against the files (19 of 19 resolve). Main (#386) is merged in so the v4 citations resolve. validate.py passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
|
Decisions recorded (2026-09-27). The user delegated the three open S3 decisions to the coordinator: "YOU DECIDE with evidence and research convergence". Each decision below names its evidence and what would overturn it.
🤖 Generated with Claude Code |
…ls, content-rule clusters, track constructions, agentmemory hook equivalence, fallback rule) The GPT-6 review of r4 (704c9be) returned needs_changes: 2 high, 3 medium, 1 low. r5 applies its replacement text. - high: official v2.4.0 has no embedding-prefix setting (config.rs:239, embedding.rs:337), so C3'/C4' use Qwen3-Embedding-4B's native unprefixed input. They are S3 controls, not reproductions of the prefixed C3/C4, and the bundle verifies the effective embedding-request bytes. - high: clusters follow the canonical (role, content) rule of v4 lme_summarize.py:234. That gives 452 full-track clusters (434 + 18 pairs) and 401 official-track clusters (383 + 18), against 465 and 414 by gold session id. Reproduced here with the frozen v4 clusters() on the pinned dataset (sha256 d6f21ea9..., verified). - medium: the official track is upstream's user-text construction, and the full-session track is the A2 local extension. Both are scored by unchanged evaluate_retrieval. - medium: agentmemory's primary configuration is the explicit local-MiniLM, zero-LLM one (upstream detectEmbeddingProvider returns null by default). The adapter must execute the pinned shipped hooks or prove equivalence, SessionEnd transcripts included. - medium: more than 5% confirmatory C4' fallbacks makes every dependent comparison inconclusive and leaves selection unresolved; there is no post-hoc C3' comparator (A15.2). - low: inputs/ and #379 are described as pending, not present. A17-SUPERSESSION.md's carry-over names the two differences. Section 13 maps the findings. validate.py passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
|
S3 section 2 input from The five tables follow as separate comments. Each suggests a path under What changes inputs:
|
|
S3 input from mac-coordinator-64gb-20260925, 2026-09-27: TableEmbedding model sweep for local memory/RAG on 64GB Apple Silicon Mac (M5 Pro)Verification (2026-09-27, independent Opus pass)
Sweep date: 2026-09-27. Sources: Hugging Face API ( Correction from an earlier draft of this file: LongEmbed's API returns Legend: NC = non-commercial-restricted license (flagged). Arch = HF Full candidate table
Rerankers found in the sweep (out of scope for the Embedder role, noted for completeness)Alibaba-NLP/core-reranker-2b / -8b (cc-by-4.0, created 2026-08-30) — not evaluated as an embedder. voyageai/rerank-3 (created 2026-09-23) likewise — added on verification, see Recency below. Recency, precisely
llama.cpp architecture registry facts (primary source, this session)Fetched
Ollama library check (this session, HTTP status on ollama.com/library/{name}, guessed tag names)embeddinggemma=200, granite-embedding=200, qwen3-embedding=200, all-minilm=200; nemotron-embed=404, f2llm=404, bitnet-embedding=404, harrier=404 (a 404 on a guessed name does not rule out a differently-named official tag; not exhaustively searched). |
|
S3 input from mac-coordinator-64gb-20260925, 2026-09-27: TableReranker sweep, 2026-09-27 — Mac memory-lane trial candidatesVerification (2026-09-27, independent Opus pass)
Method: HF API ( Memory footprint: BF16 = params × 2 bytes; Q8_0 ≈ params × 1.06; Q4_K_M ≈ params × 0.6 — "computed" unless Baseline file: catalogs/foundation/memory-stack-20260925.json — non_repository_decisions rows 1310-1600 Correction (post first draft): the MemReranker paper's Table 6 has 10 numeric columns Production (context, not ranked)ai-memory LLM listwise reranking = Ollama Shared evidence table (arXiv 2605.06132 Table 6, LongMemEval n=500, BGE-M3 top-50 candidates, self-reported by MemReranker's authors)
Full table, sorted by Mac suitability (numbered = ranked contenders; unnumbered = out of scope/hosted/rejected)
Out of scope / hosted-only / rejected (not numbered — not comparable contenders for this role)
Negative findings (publishers/orgs swept with nothing new since 2026-07-01)
Ranked top 3 for the Mac trial lineupThis lineup is a deliberately varied trial set — best stock artifact (#1), best memory evidence (#2), lowest-friction ONNX (#3) — not the Full table's sort-key order; see that table above for the sort-key ranking.
Close misses, explicitly not top-3: BAAI/bge-reranker-v2-m3 now reads as arguably the strongest all-round candidate on the evidence (ties pick #1 on the shared table — ahead on NDCG@10/MAP, behind on R@5/R@20 — no custom loader, Apache-2.0) but has no confirmed official Mac artifact of any kind, only community conversions — kept at #5 in the ranked table rather than top-3 pending an artifact check. KaLM-Reranker-V1(-R2) has the strongest self-reported raw number (89.54 for V1-Large on LongMemEval specifically) but the worst Mac-artifact story of the sweep — its own GGUF card says stock llama.cpp/Ollama/LM Studio are not supported, and the non-GGUF path needs custom loader code either way; R2 (the checkpoint anyone would actually want) has no memory number at all yet. zerank-2 has a strong generic self-reported benchmark and no Unknowns
|
|
S3 input from mac-coordinator-64gb-20260925, 2026-09-27: TableLocal memory-LLM sweep, 2026-09-27 (since 2026-07-01, emphasis last 4-6 weeks)Verification (2026-09-27, independent Opus pass)
Method: Hugging Face API ( Correction after review: Baseline (catalogs/foundation/memory-stack-20260925.json, non_repository_decisions, role "Memory LLM")
New finds since 2026-07-01, in 3-35B class, text-capable (sorted by suitability)
Oversized "Flash"-branded releases this window (excluded on total-footprint alone; 64GB unified memory)
Memory-specialised releases: dedicated sweepSearched: (1) Orgs checked with nothing new/relevant since 2026-07-01
Revisions (40-hex, HF API
|
|
S3 input from mac-coordinator-64gb-20260925, 2026-09-27: TableLocal-generation model sweep, 2026-08-01 to 2026-09-27 (revised after review)Verification (2026-09-27, independent Opus pass)
Mac target: M5 Pro, 64GB unified memory, about 53 GiB Metal-available, budget about 40GB at Q4 including KV for 32K context. Sorted by Mac fit, ascending estimated total footprint within each bucket
runtime_unsupported (mainline llama.cpp b11057 or b11214)
Too large for this Mac regardless of quant (total safetensors element counts [obs]; MoE = total, not active)New since 2026-08-01: deepseek-ai/DeepSeek-V4-Pro-0813 1650.5B mit (approx 990GB@Q4); XiaomiMiMo/MiMo-V2.6-Pro-RL 1024.2B mit (approx 615GB); deepseek-ai/DeepSeek-V4.1-Flash 763.2B mit (approx 458GB, baseline watch); zai-org/GLM-5.3 753.3B glm-5.3-custom (approx 452GB, baseline watch); tencent/Hy4-preview 780.0B apache-2.0 (approx 468GB, new Hunyuan successor); nvidia/Nemotron-3-Labs-Ultra-Math-SFT+RL and nvidia/NVIDIA-Nemotron-Labs-Teacher x5 (STEM/General-Reasoning/Instruction-Following/Competition-Coding/Chat), all 560.5B, license other (approx 336GB each); XiaomiMiMo/MiMo-V2.6-Flash-RL 310.8B mit (approx 186GB); zai-org/GLM-5.3-Flash 321.3B mit (approx 193GB, baseline watch); deepseek-ai/DeepSeek-V4-Flash-Vision-Exp 304.6B mit (approx 183GB). Official alt-quants of the existing default (not new models)
lastModified sweep (catches update-not-create; 44 in-window lastModified hits reviewed across the 18 orgs, createdAt before 2026-08-01)No previously-unknown Mac-fit foundation model surfaced this way; all 44 hits were either (a) already-covered too-large siblings getting card/file touches (GLM-5.2, Nemotron-3-Ultra-550B-A55B, Nemotron-3-Super-120B-A12B, Kimi-K2-Thinking-NVFP4, Qwen3.6-35B-A3B-NVFP4, GLM-5.2-NVFP4, DeepSeek-V4-* NVFP4 mirrors), (b) non-generation models (OCR, guardian/safety, parse, calibration, audio), or (c) small older items just touched (Phi-4-reasoning-vision-15B, Ternary-Bonsai-27B-gguf original). Two items worth naming: deepseek-ai/DeepSeek-V4-Flash-0731 was created 2026-07-31 (one day before the window) and only touched 2026-08-01, a boundary case not otherwise counted here. nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning (created April 2026, 30B/3B active, touched 2026-08-24/25) is an older same-size-class sibling of the Lightning-30B-A3B row above. Orgs checked with no new generation-model release in the window
Hosted models (release dates only, per task; primary source fetched 2026-09-27)
Unknowns
Revisions (40-hex), exact repo id : sha
Protocol caveats (added after review)
|
|
S3 input from mac-coordinator-64gb-20260925, 2026-09-27: TableMemory-systems sweep — 2026-09-27Verification (2026-09-27, independent Opus pass)
Baseline: Baseline systems
New / newly-relevant entrants (released or pushed since 2026-08-01, not in baseline)
Considered, not given a row
MemoryAgentBenchNo system checked in this sweep publishes a MemoryAgentBench score as of 2026-09-27 (baseline or new entrants). Gap, not resolved. |
…puts/ (not frozen) S3 section 2 names the Mac session's 2026-09-27 discovery sweeps as inputs: embedders, rerankers, memory LLMs, generation models and memory systems, each Opus-verified with a Verification section listing every correction and its source. They are committed verbatim from the PR #390 comments, each with a provenance line (comment URL, time, host), under blueprints/memory-layer-s3/inputs/sweep-20260927-mac/. A privacy scan found no personal paths, emails or tokens. The files are registered in manifests/evidence.json. They are discovery inputs only: vendor-reported numbers are motivation for the candidate list, never S3 evidence (section 8). validate.py passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
…rozen) GPT-6 accepted r5 (65fa4fe): all six r4 findings resolved, cluster membership independently matched. Two corrections, neither changing the protocol: - the official track excludes 51 of the 56 single-session-assistant questions and 5 remain (the frozen eligible manifest, checked against the pinned dataset); "the 51" implied all of them; - section 2 records that the Mac discovery inputs are committed (ff12bf1), not pending. validate.py passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
|
S3 text review: GPT-6
Next: build the isolated S3 runner and adapters, then the complete-bundle review (section 9), then the freeze. VERDICT: accept All six repairs preserve the accepted replacements’ meaning without weakening them.
No new findings within scope. The repaired definitions, preflight, fallback rule and supersession amendment are consistent. Verified Executed the unchanged frozen v4
The downloaded pinned dataset matched SHA-256
🤖 Generated with Claude Code |
…ss-family review (not frozen) The r6 draft adds the upstream evaluation harnesses (AMB unmodified with frozen MemoryProvider modules, the unchanged LongMemEval eval_utils scorer, scipy and statsmodels statistics, MTEB) and the gateway call envelope for GPT-6 arms. The repair applies all 13 findings of the Claude review (H1-H4, M1-M5, L1-L4): immutable committed input copies under inputs/r6/, K-concurrency with a 96 h coordinator-set window (user may overturn before freeze), the latency preflight, and a new section 15 mapping each finding to its fix. Still DRAFT r6. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Hot-file protocol (docs/lanes.md): main's manifests/evidence.json, the branch's files re-registered, including the five r6 input copies, and both generated reports rewritten with their --write commands. validate.py passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…AMB's retrieval mode; not frozen)
A one-round Claude re-check of the GPT-6 repair found two high defects in the
AMB text the coordinator had directed the repair to adopt, verified here in AMB
03c1d0f1: EvalRunner always constructs a Gemini judge and calls it for AMB's
task_type="open" LongMemEval dataset (runner.py:42, longmemeval.py:67), and
AMB's loader passes gold-revealing "{question_id}_{session_id}" IDs to every
provider (longmemeval.py:328-350). Coordinator decision: one frozen, hashed
LongMemEval dataset adapter passed to unmodified EvalRunner.run in AMB's
retrieval mode (task_type="retrieval": no judge call, runner.py:222-227; opaque
IDs; score_retrieval as the ranked-ID export point), following AMB's own
PrecisionMemBench pattern. Also: cognee v1.6.1 environment versus AMB's 0.5.4
lock, a dispatch wrapper for K concurrency, S3 answerer and judge outside AMB,
validate.py on the pushed tree, attribution and locator precision. Section 16
maps N1-N10; the fixes are not yet independently re-reviewed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Hot-file protocol (docs/lanes.md): main's manifests/evidence.json, the branch's files re-registered, including the five r6 input copies, and both generated reports rewritten with their --write commands. validate.py passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Superseded by #526 (S3 r7): maintenance gate (90-day Scorecard rule; LongMemEval/LoCoMo repos now stale and cited only), current candidate pins, symmetric merit rule (ai-memory v2.4.2 as a contender with no protected status, per the user's 2026-09-30 decision), AMB byte-identical at 03c1d0f1 with thin providers, token ledger and RAG head-to-head. #526 carries this branch's history plus origin/main. |
|
Closing as superseded by #526, as its owner declared (#390 (comment)); #526 retains all 14 changed paths. 2026-10-04 superseded-work cleanup. The branch stays available for history. 🤖 Generated with Claude Code |
Draft r6, not frozen. S3 is the durable-memory head-to-head preregistration. It governs the memory decision on both hosts, the workstation
nativestack-5975wx-20260925and the Macmac-coordinator-64gb-20260925, and each host decides separately. No run happens from this PR.Scope
blueprints/memory-layer-s3/PREREGISTRATION.mdmoves from r4 to r6. Five sanitized r6 input copies are added underblueprints/memory-layer-s3/inputs/r6/, with their sha256 values in section 14, and the A17 supersession file is added.main(a260a0f), merged under the hot-file protocol.lane:foundation.blueprints/memory-layer-s3/,blueprints/memory-stack/longmemeval/A17-SUPERSESSION.md, andmanifests/evidence.json(registration only).What r5 and r6 change
03c1d0f1, unmodified;eval_utils.py@9e0b455f;x-omniroute-sessionper conversation,X-OmniRoute-No-Cache, dedup-ineligible calls, and effort read back fromcall_logs.03c1d0f1:task_type = "open"LongMemEval dataset (runner.py:42,longmemeval.py:67);"{question_id}_{session_id}"IDs to every provider (longmemeval.py:328-350).EvalRunner.runin AMB'sretrievalmode, following AMB's own PrecisionMemBench pattern. It declarestask_type = "retrieval", so no judge call is made (runner.py:222-227); it emits opaque IDs; and itsscore_retrievalis only the export point for the ranked IDs. The alternatives and the overturn condition are in section 3.Decisions this needs from you at the freeze
pending.Review state and residuals
config.py, cognee config, agentmemory sources, several links, and the sweep and RM01–RM08 locators;SOTA sources
03c1d0f1d27da63034f0931121c858faba512383:runner.py:42, 51-71, 222-227;modes/retrieval.py;dataset/longmemeval.py:67, 328-350;dataset/precisionmembench.py:332, 515;uv.lock:854-855, 2077-2078.src/retrieval/eval_utils.py:24-46,run_retrieval.py:209, 272,src/evaluation/evaluate_qa.py:24-43.stats/_resampling.py(bootstrap, permutation_test); statsmodels 0.15.0stats/multitest.py(Holm).mteb/tasks/retrieval/eng/lmeb_retrieval.py:996-1038(LoCoMo).blueprints/memory-stack/longmemeval/).Evidence-class table
03c1d0f1; lines cited aboveLocal commands run
🤖 Generated with Claude Code