Skip to content

Memory layer S3: head-to-head preregistration draft r6 (both hosts; upstream AMB harness in retrieval mode; not frozen) - #390

Closed
seathatflowsinourveins wants to merge 11 commits into
mainfrom
claude/memory-layer-s3-prereg-20260927
Closed

seathatflowsinourveins wants to merge 11 commits into
mainfrom
claude/memory-layer-s3-prereg-20260927

Conversation

@seathatflowsinourveins

@seathatflowsinourveins seathatflowsinourveins commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Draft r6, not frozen. S3 is the durable-memory head-to-head preregistration. It governs the memory decision on both hosts, the workstation nativestack-5975wx-20260925 and the Mac mac-coordinator-64gb-20260925, and each host decides separately. No run happens from this PR.

Scope

  • What changes: blueprints/memory-layer-s3/PREREGISTRATION.md moves from r4 to r6. Five sanitized r6 input copies are added under blueprints/memory-layer-s3/inputs/r6/, with their sha256 values in section 14, and the A17 supersession file is added.
  • Base: current main (a260a0f), merged under the hot-file protocol.
  • Lane: lane:foundation.
  • Owned paths touched: blueprints/memory-layer-s3/, blueprints/memory-stack/longmemeval/A17-SUPERSESSION.md, and manifests/evidence.json (registration only).

What r5 and r6 change

  • r5: GPT-6's replacement text for its r4 review (2 high, 3 medium, 1 low; section 13).
  • r6, the user's 2026-09-27 directions:
    • GPT-6 through each host's OmniRoute for every LLM role;
    • isolated upstream deployments of cognee and Hindsight;
    • current local retrieval challengers, selected on MTEB 2.21.8 LoCoMo, never on LongMemEval;
    • research-memory usefulness;
    • early lifecycle checks.
  • r6 evaluation harness, upstream rather than self-written:
    • AMB 03c1d0f1, unmodified;
    • unchanged LongMemEval eval_utils.py @9e0b455f;
    • SciPy 1.18.1 and statsmodels 0.15.0 for the statistics;
    • MTEB 2.21.8.
  • r6 gateway call envelope: one x-omniroute-session per conversation, X-OmniRoute-No-Cache, dedup-ineligible calls, and effort read back from call_logs.
  • Repair round (GPT-6): all 13 findings of a Claude cross-family review (4 high, 5 medium, 4 low), mapped in section 15.
  • Re-check round (Claude, one round): 2 high, 4 medium and 4 low defects (N1–N10), mapped in section 16. The two high ones sat in the AMB text the coordinator had told the repair to adopt, and both were verified in AMB 03c1d0f1:
    • AMB always constructs a Gemini judge and calls it for its task_type = "open" LongMemEval dataset (runner.py:42, longmemeval.py:67);
    • AMB's loader passes gold-revealing "{question_id}_{session_id}" IDs to every provider (longmemeval.py:328-350).
  • Coordinator decision: one frozen, hashed LongMemEval dataset adapter, passed to unmodified EvalRunner.run in AMB's retrieval mode, following AMB's own PrecisionMemBench pattern. It declares task_type = "retrieval", so no judge call is made (runner.py:222-227); it emits opaque IDs; and its score_retrieval is only the export point for the ranked IDs. The alternatives and the overturn condition are in section 3.

Decisions this needs from you at the freeze

  1. Confirm the A17 supersession.
  2. Approve a workstation GPU window, if a selected local embedder or reranker needs one.
  3. Apple Container on the Mac for Hindsight, or leave Mac Hindsight pending.
  4. New: the 96 h per-host wall-clock window for the confirmatory retrieval run is the coordinator's decision, and you may overturn it before the freeze. K concurrency is a throughput parameter set on the development split, never a quota throttle.

Review state and residuals

  • Reviews so far: r1 through r5 had GPT-6 review rounds (sections 10–13). r6 had one Claude cross-family review, one GPT-6 repair round and one Claude re-check. The coordinator applied the re-check fixes, and nobody has independently re-reviewed them yet.
  • Residuals, all for the section 9 complete-bundle review (GPT-6):
    • the N1–N10 fixes;
    • source files the re-check could not see: OmniRoute idempotency and dedup files, Hindsight config.py, cognee config, agentmemory sources, several links, and the sweep and RM01–RM08 locators;
    • the adapter, the dispatch wrapper and the provider modules. These are not written yet; the bundle review checks them before the freeze.

SOTA sources

Evidence-class table

Claim Evidence class Command / receipt
AMB judge, ID and lock facts behind N1 to N3 source_review AMB clone at 03c1d0f1; lines cited above
The r6 input copies equal their sanitized sources local_integration byte comparison after sanitization; the sha256 values are in section 14
Every S3 result not run none; this PR is a preregistration draft

Local commands run

$ python3 scripts/validate.py
{"components": 69, "hashed_files": 7361, "profiles": 4, "receipts": 159, "status": "passed"}

🤖 Generated with Claude Code

Scout and others added 2 commits September 27, 2026 00:48
The merit-only head-to-head that decides each host's durable-memory
configuration: ai-memory v2.4.0 as the reference arm, 7 challengers, native
Claude and Codex memory controls, and frozen reserves. r2 answers all 11
findings of the GPT-6 review of r1: eligibility applies to every
configuration, superiority and non-inferiority-plus-cost routes, a
qmax-minus-epsilon cost selection, cluster bootstrap statistics, a full
dataset and adapter contract, two token ledgers, a defined latency statistic,
usefulness lanes, and per-host qualification. Not frozen: it freezes only
after a GPT-6 accept on the complete bundle.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
…cement text (not frozen)

r2 resolved 9 of r1's 11 findings. r3 applies GPT-6's replacement text for the
2 partly resolved findings and the 4 new ones:
- a usefulness floor applied before either replacement route (new high);
- a one-sided 5th-percentile lower bound;
- one Holm family holding both the superiority and the non-inferiority
  hypotheses;
- a single frozen complete-workload decision ledger, where payload-only counts
  support only a payload-size claim;
- a frozen reference-availability preflight, cluster manifest, seed
  aggregation, question weighting and RNG.
The next GPT-6 review is the bundle review before the freeze.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Sep 27, 2026
Scout and others added 2 commits September 27, 2026 09:39
….4.0 reranker feasibility; v4-based runner rules

S3 draft r4 (not frozen):
- Reference feasibility from source. Official ai-memory v2.4.0 ships the
  LLM reranker: AI_MEMORY_RERANKER=llm in
  crates/ai-memory-cli/src/config.rs:413-420,1744-1759,
  crates/ai-memory-llm/src/reranker.rs and docs/llm-providers.md:185-192.
  The docs name no reranker model, so the preregistered qwen3.5-9b-64k
  stays. Its base is pinned by v4 pins.json:141-147, and the runner verifies
  the base, the effective template and parameters, and the served blob
  hashes.
- C3' and C4' are defined from the v4 driver's C3/C4. agentmemory's primary
  arm is its shipped hook path, the A-protocol's D2h shape, with its v4 pins.
- Late-candidate rule: later sweep survivors enter only through a dated
  amendment arm, like reserves admitted after confirmatory execution. The
  32-layer sweep is held for the user's Gates A and B.
- D-mac-prod: the Mac's production 19b6429 in a C4-shaped configuration, a
  labelled diagnostic outside the decision family, at the Mac session's
  request. The 19b6429 / v2.4.0 / v2.4.1 lineage is recorded.
- Harness: v4 (#386) is the frozen adapter source, and the S3 runner is new
  code with the isolation rules from the #386 GPT-6 reviews. The cache
  proxy's omission is preregistered from A16.3. Only the public reranker
  Modelfile and its base pin are needed.
- A workstation GPU window, a scheduled production stop the user approves;
  per-host qualification, including reranker timeout latency; Mac Hindsight
  pending the user's Apple Container decision.
- S3's statistics govern, and the A-protocol's do not transfer. Section 12
  covers the relationship to A1-A16.3.

New blueprints/memory-stack/longmemeval/A17-SUPERSESSION.md (PROPOSED): R1
is not run under A1-A16.3, and S3 supersedes it on both hosts. No
confirmatory arm has run on any host, so no look is spent. It takes effect
only when the user confirms it at the S3 freeze. No session held an A17
draft (checked with the sessions that had owned it).

Every path:line citation in the new text was checked against the files
(19 of 19 resolve). Main (#386) is merged in so the v4 citations resolve.
validate.py passed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
@seathatflowsinourveins seathatflowsinourveins changed the title Memory layer S3: head-to-head preregistration draft r3 (not frozen; freezes after the sweep's durable-memory candidates and a GPT-6 bundle accept) Memory layer S3: head-to-head preregistration draft r4 (governs both hosts; A17 supersession proposed; not frozen) Sep 27, 2026
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Decisions recorded (2026-09-27). The user delegated the three open S3 decisions to the coordinator: "YOU DECIDE with evidence and research convergence". Each decision below names its evidence and what would overturn it.

  1. A17 → S3 supersession: confirmed. It takes effect with the S3 freeze (section 9), which still requires a GPT-6 accept on the complete bundle.
    • Evidence: no confirmatory arm of A1–A16.3 has run on any host (retire-vela-velanext.md:43-44; memory-stack-20260925.json:1318; mac-single-writer-staged.md:48), so no look is spent. One governing protocol replaces two overlapping ones with different statistics. No session held an A17 draft.
    • Overturned by: a record of any confirmatory A1–A16.3 result on any host, found before the freeze.
  2. Workstation GPU window: approved. It runs as a scheduled, ownership-checked stop and restore of production vLLM (127.0.0.1:18231) and llama.cpp (127.0.0.1:18232), with a restore timer and a /health check after restore, only when the S3 runs start. The runner never stops a unit on its own.
    • Evidence: the RTX 4090 is fully held by those two services (host use receipts), and the v4 driver's own gpu_idle gate refuses a busy GPU.
    • Overturned by: a measured co-residency that fits every S3 arm alongside production without changing results.
  3. Mac Hindsight: Apple Container. It is the macOS platform profile's own Postgres selection. Relinking the pg0-embedded Postgres would modify the upstream artifact, so it would not be a faithful arm.
    • The owner's sudo install is an operating-system consent step: the Mac session asks the user for it when it is ready to run the arm. Until then Mac Hindsight is pending, not refuted.
    • Overturned by: an upstream pg0-embedded release whose macOS Postgres runs without Homebrew.

🤖 Generated with Claude Code

https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5

…ls, content-rule clusters, track constructions, agentmemory hook equivalence, fallback rule)

The GPT-6 review of r4 (704c9be) returned needs_changes: 2 high, 3 medium,
1 low. r5 applies its replacement text.
- high: official v2.4.0 has no embedding-prefix setting (config.rs:239,
  embedding.rs:337), so C3'/C4' use Qwen3-Embedding-4B's native unprefixed
  input. They are S3 controls, not reproductions of the prefixed C3/C4, and
  the bundle verifies the effective embedding-request bytes.
- high: clusters follow the canonical (role, content) rule of v4
  lme_summarize.py:234. That gives 452 full-track clusters (434 + 18 pairs)
  and 401 official-track clusters (383 + 18), against 465 and 414 by gold
  session id. Reproduced here with the frozen v4 clusters() on the pinned
  dataset (sha256 d6f21ea9..., verified).
- medium: the official track is upstream's user-text construction, and the
  full-session track is the A2 local extension. Both are scored by
  unchanged evaluate_retrieval.
- medium: agentmemory's primary configuration is the explicit
  local-MiniLM, zero-LLM one (upstream detectEmbeddingProvider returns null
  by default). The adapter must execute the pinned shipped hooks or prove
  equivalence, SessionEnd transcripts included.
- medium: more than 5% confirmatory C4' fallbacks makes every dependent
  comparison inconclusive and leaves selection unresolved; there is no
  post-hoc C3' comparator (A15.2).
- low: inputs/ and #379 are described as pending, not present.
A17-SUPERSESSION.md's carry-over names the two differences. Section 13
maps the findings. validate.py passed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

S3 section 2 input from mac-coordinator-64gb-20260925 (2026-09-27): five verified sweep tables

The five tables follow as separate comments. Each suggests a path under inputs/sweep-20260927-mac/. Every table came from a Sonnet sweep against primary sources: the Hugging Face API, the model cards and LICENSE files, the llama.cpp source at b11146 and b11214, registry JSON, and the GitHub releases API. An independent Opus pass then checked the claims that decide each ranking, and the tables were corrected. Each table's "Verification" section lists the corrections and their sources. Per the S3 merit rules, the published and self-reported scores here (LMEB, AA-LCR, vendor cards, and the MemReranker paper's Table 6) are discovery evidence, not S3 evidence.

What changes inputs:

  • Memory LLM (consolidation/rerank). The AA-LCR v1.1 scores are recoverable from Artificial Analysis's page data. The mapping reproduces the baseline's own figures exactly (Qwen3.5-9B 46.0/70.0, Qwen3.6-35B-A3B 64.3/71.7, Nemotron 60.3).
  • Gemma-4 license. license_link resolves to plain Apache-2.0.
  • LFM2.5-8B-A1B. It has an Ollama tag, lfm2.5:8b. Its LFM Open License v1.0 withholds commercial use from entities with US$10M or more in annual revenue.
  • Embedders. No new dense text embedder in the last 4-6 weeks.
    • Nemotron-3-Embed-1B beats F2LLM-v2-4B on every LMEB figure, but has no llama.cpp or Mac path (Ministral3Model is unregistered at b11146).
    • F2LLM-v2-4B leads only for a llama.cpp deployment: Apache-2.0, Qwen3Model registered, the same path as production.
    • "LMEB rank" is the leaderboard's Borda rank.
  • Rerankers. MemReranker-4B has the best LongMemEval NDCG@10/MAP (0.8354/0.8043, self-reported, community MLX only). Qwen3-Reranker-0.6B is the only one with a maintainer-demonstrated stock GGUF (b11057 untested). BGE-v2-m3 vs Qwen3-Reranker-0.6B is effectively a tie.
  • ai-memory lineage. 19b6429 is 46 commits ahead of and 62 behind v2.4.1 (merge base 8ee81bed), and v2.4.0 is its ancestor. This goes into the reference-arm rationale.
  • Generation. No in-window model has task-level evidence against Qwen3.8-27B.
    • granite-4.2-30b: official GGUF, MLX and Ollama, about 27.8 GB at 32K.
    • Nemotron-3.5-Lightning-30B-A3B: about 27.0 GB at 32K, KV 0.20 GB.
    • Either one would need the li26 protocol.
  • Memory systems. Trial pins unchanged: agentmemory 0.9.29, MemPalace 3.10.0, Hindsight 0.10.1. For Hindsight's capture path, pin integrations/coding-agents/v0.7.0.
    • Mac Hindsight is pending (Apple Container, the owner's install).
    • claude-mem's interactive installer pre-selects hosted CMEM Pro.
    • GBrain's hang issues #5284 and #5449 are still open.
    • Memorix v1.9.6 and mcp-memory-service v11.14.0 are Apache-2.0 defer candidates.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

S3 input from mac-coordinator-64gb-20260925, 2026-09-27: embed sweep table (suggested path: inputs/sweep-20260927-mac/embed.md). Sonnet sweep, independently verified by an Opus pass against primary sources; its "Verification" section lists every correction applied. The published and self-reported numbers here are discovery evidence, not S3 evidence.

Table

Embedding model sweep for local memory/RAG on 64GB Apple Silicon Mac (M5 Pro)

Verification (2026-09-27, independent Opus pass)

  • nvidia/Nemotron-3-Embed-1B-BF16's base model corrected to mistralai/Ministral-3-3B-Instruct-2512 (not the 8B sibling). Source: embed/raw/nvidia.json (base_model:mistralai/Ministral-3-3B-Instruct-2512 / base_model:finetune:mistralai/Ministral-3-3B-Instruct-2512 tags on the 1B repo, distinct from the 8B repo's own Ministral-3-8B-Instruct-2512 tags).
  • Qwen3-Embedding-4B's LongMemEval 50.21 is 93rd of 105 on LMEB, not the weakest; the weakest is sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 at 28.58. "LMEB rank" is redefined below as the leaderboard's own Borda rank over 22 tasks (not a re-sort by the displayed mean), and each of the four picks now also states its position by mean and by LongMemEval. Source: embed/lmeb_all105.txt (105-row leaderboard dump; positions independently recomputed by sorting this file by its mean and LongMemEval columns: F2LLM-v2-4B 6th/5th, Nemotron-1B 2nd/3rd, harrier-0.6b 8th/46th, Qwen3-Embedding-4B 51st/93rd — all four match exactly).
  • This file has no dedicated ranking/rationale section, so the Nemotron-1B vs. F2LLM-v2-4B comparison and the llama.cpp-vs-PyTorch/MPS framing are stated here rather than invented as a new section: on every LMEB figure (Borda rank 3 vs 10, mean 61.50 vs 58.68, LoCoMo 50.77 vs 43.32, LongMemEval 81.58 vs 81.42) Nemotron-1B beats F2LLM-v2-4B outright. F2LLM-v2-4B's edge is deployment-shaped, not evidence-shaped: it is Qwen3Model (registered in llama.cpp's conversion/qwen.py:159, apache-2.0, same size/path as production's own Qwen3-Embedding-4B row), whereas Nemotron-1B's bare Ministral3Model is not registered there (see the architecture-registry facts below) — so for a llama.cpp deployment target F2LLM-v2-4B is the pragmatic Fix native worker timeout cleanup on macOS #1, but for a PyTorch/MPS target the evidence favors Nemotron-1B on every figure. harrier-oss-v1-0.6b's LongMemEval position is only 46th (of 105), well behind both.
  • Added repos created in the window that are not text embedders, so the "no new dense text embedder" conclusion stands unchanged: tencent/EVIE-Preview-4.5B, tencent/EVIE-4.5B, tencent/EVIE-8B (image-document retrievers), tencent/WeVisDoc-2B/4B (image-document retrievers), voyageai/rerank-3 (a reranker). None of these appear in embed/raw/tencent.json or embed/raw/voyageai.json (the original sweep's own org listings), confirming they were not in scope for this sweep's HF org pulls; added here on the independent pass. Source: https://huggingface.co/tencent/EVIE-Preview-4.5B, https://huggingface.co/tencent/WeVisDoc-2B, https://huggingface.co/voyageai/rerank-3.

Sweep date: 2026-09-27. Sources: Hugging Face API (/api/models, /api/models/{org}/{name}, per-model raw/main/config.json, README.md, LICENSE), MTEB/RTEB leaderboard backend (https://mteb-leaderboard-backend.hf.space/v1/benchmarks/{LMEB,RTEB(eng, beta),LongEmbed}/scores, snapshot this session), ggml-org/llama.cpp conversion/*.py at commit 59657a613ab0fa4ab327d6c790123dff30bfbd67 (catalog tag b11057) and 7fe450e19305b828c199d602c23a8337aaa1f03b (v0.5.0/b11146) — convert_hf_to_gguf.py is byte-identical between the two, ollama.com/library/{name} HTTP status, gh api repos/embeddings-benchmark/results/commits for submitter attribution. Baseline: catalogs/foundation/memory-stack-20260925.json (non_repository_decisions, role Embedder, lines 937-1309) and catalogs/us-equities/models.md (lines 19-25).

Correction from an earlier draft of this file: LongEmbed's API returns meanTask: null (not 0) for any model missing one of its 6 sub-tasks. A first pass rendered null as 0.00 for Qwen3-Embedding-4B, embeddinggemma-300m, harrier-oss-v1-0.6b/27b and others — this looked like "fails long-document retrieval" but is actually "only the synthetic Passkey-Retrieval sub-task is scored; the other 5 are not evaluated for this model, so no overall mean exists." Verified directly against longembed.json scoresByTask (raw, not the derived mean) for each row below. Do not read a blank LongEmbed mean as a zero.

Legend: NC = non-commercial-restricted license (flagged). Arch = HF config.json architectures[0], and whether that exact class string is registered in llama.cpp's conversion/ package. LMEB evidence = leaderboard-hosted (KaLM/MTEB embeddings-benchmark/results repo), zero-shot (no train-set contamination flag) unless noted; submitter accounts checked via gh api .../commits show no visible vendor affiliation for either Nemotron or F2LLM-v2 (same maintainer accounts, e.g. ItsukiFujii, merge both) — treat as leaderboard/community-submitted, distinct from a vendor's own blog claim (DV). RTEB(eng,beta): meanTask/meanPublic/meanPrivate are null for 264/274 rows (a private-holdout design); only rank (of 274) is usable there. LMEB rank = the leaderboard's own Borda rank over its 22 constituent tasks, not a re-sort by the displayed mean column — the two orders diverge (e.g. Nemotron-1B is Borda rank 3 but 2nd by mean; F2LLM-v2-4B is Borda rank 10 but 6th by mean); each pick below states its position by mean and by LongMemEval alongside its Borda rank for that reason.

Full candidate table

Model (HF repo) rev (sha, short) Release/lastModified Params Dim Max tok License Mac artifact llama.cpp arch (b11057+) Mem Q8/BF16 Key evidence (source)
google/embeddinggemma-300m (current Mac pin) 57c266a7 created 2025-07-17, lastMod 2025-09-25 302.9M 768 (MRL to 128) 2048 gemma (gated=manual, ToS not plain OSS) Official: ggml-org/embeddinggemma-300M-GGUF, -qat-q8_0-GGUF, -qat-q4_0-GGUF; Ollama embeddinggemma (HTTP 200) Gemma3TextModel — registered (conversion/gemma.py:182) 0.30/0.61 GB No LMEB row (LongMemEval/LoCoMo: unknown). LongEmbed: only Passkey-Retrieval scored (60.75%); other 5 long-doc tasks not evaluated, no overall mean (not "fails," genuinely unverified)
microsoft/bitnet-embedding-0.6b 16176d10 created/lastMod 2026-07-15/17 596M (1.58-bit native) 1024 32768 MIT Official GGUF bitnet-embeddings-0.6b-bf16-i2_s.gguf shipped pre-quantized by publisher GGUF-only repo, no config.json; base is BitNet b1.58 — BitnetForCausalLM/BitNetForCausalLM registered (conversion/bitnet.py:11); i2_s is a native llama.cpp quant type ~0.12GB native 1.58-bit (Q8-equiv 0.60/BF16-equiv 1.19) DV only (publisher README table, MTEB v2 mean 67.49 — below harrier-oss-v1-0.6b's 69.0 at the same size). No LMEB/RTEB/LongEmbed row → LongMemEval/LoCoMo unknown
microsoft/bitnet-embedding-270m 5f1c2fd1 created/lastMod 2026-07-15/17 268M (1.58-bit) 640 32768 MIT same as above same ~0.05GB native (Q8-equiv 0.27/BF16-equiv 0.54) DV: MTEB v2 mean 66.26. No board row
sentence-transformers/all-MiniLM-L6-v2 (agentmemory control, baseline) 1110a243 lastMod 2026-06-01 22.7M 384 256 apache-2.0 agentmemory ships Xenova/all-MiniLM-L6-v2 ONNX (community org, not sentence-transformers itself — community-only per task rule) n/a (BERT-family; not individually re-checked) 0.02/0.05 GB LMEB rank 99, LongMemEval 67.94, LoCoMo 25.86. LongEmbed: full 6-task mean 29.81 (rank 136/201, genuinely weak on long docs)
mixedbread-ai/mxbai-embed-large-v1 — lastMod 2026-01-23 335M 1024 512 apache-2.0 Official ONNX in repo (onnx tag) not individually checked 0.34/0.67 GB LMEB rank 71, LongMemEval 74.53, mean 50.24. RTEB rank 70. LongEmbed full mean 38.92 (rank 91)
LiquidAI/LFM2.5-ColBERT-350M-GGUF / LFM2-ColBERT-350M-GGUF — lastMod 2026-06-22 / 2026-06-29 353M late-interaction (multi-vector) 128000 LFM Open License v1.0 (verified via LICENSE text: free if annual revenue <$10M, Sections 1 & 5) Official GGUF from publisher Lfm2Model/Lfm2BidirectionalModel registered (conversion/lfm2.py:68) 0.35/0.71 GB Dense sibling LiquidAI/LFM2.5-Embedding-350M: LMEB rank 72, LongMemEval 62.79, mean 51.05. ColBERT variant needs late-interaction (multi-vector) retrieval infra, not a drop-in for single-vector cosine RAG
Qwen/Qwen3-Embedding-4B (production baseline) 5cf2132a lastMod 2025-06-20 (unchanged, cross-checked this session — sha matches baseline JSON exactly) 4.02B 2560 32768 apache-2.0 Official GGUF Qwen/Qwen3-Embedding-4B-GGUF; Ollama qwen3-embedding:4b df5bd2e3c74c Qwen3ForCausalLM/Qwen3Model registered (conversion/qwen.py:159) 4.02/8.04 GB (Ollama tag is Q4_K_M, smaller) LMEB rank 47 (Borda; 51st by mean, 93rd by LongMemEval), LongMemEval 50.21 (93rd of 105 by LongMemEval, not the weakest scored model on the board — that is sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 at 28.58), LoCoMo 30.08, mean 53.30. RTEB rank 194/274. LongEmbed: only Passkey-Retrieval scored (84.25%), no full mean
codefuse-ai/F2LLM-v2-4B 3e565254 created 2026-03-09, lastMod 2026-09-03 (metadata touch only — not a new release) 4.02B 2560 40960 apache-2.0 None official (safetensors only) Qwen3Model — registered, same class as Qwen3-Embedding 4.02/8.04 GB LMEB rank 10 (Borda; 6th by mean, 5th by LongMemEval), LongMemEval 81.42, LoCoMo 43.32, mean 58.68, zeroShotPct=100/trainedOnTasks=[] (no contamination flag, same as Nemotron). RTEB rank 23/274. LongEmbed full 6-task mean 75.36 (rank 7/201)
codefuse-ai/F2LLM-v2-1.7B c5650fe0 2026-03-09 / 2026-09-03 1.72B 2048 40960 apache-2.0 none official Qwen3Model — registered 1.72/3.44 GB LMEB rank 14, LME 78.82, LoCoMo 45.65, mean 57.60. LongEmbed full mean 71.43 (rank 11)
codefuse-ai/F2LLM-v2-8B 2f1beedb 2026-03-09 / 2026-09-03 7.57B 4096 40960 apache-2.0 none official Qwen3Model — registered 7.57/15.14 GB LMEB rank 8, LME 81.46, LoCoMo 42.29, mean 58.64. RTEB rank 20. LongEmbed full mean 76.64 (rank 5, best of the family)
codefuse-ai/F2LLM-v2-0.6B 8786315a 2026-03-02 / 2026-09-03 596M 1024 40960 apache-2.0 Official ONNX (onnx/model.onnx shipped by codefuse-ai itself — counts as official per task rule) Qwen3Model — registered 0.60/1.19 GB LMEB rank 37, LME 76.22, LoCoMo 39.54, mean 55.78. LongEmbed full mean 70.81 (rank 21)
codefuse-ai/F2LLM-v2-14B — 2026-03-09 / 2026-09-03 13.99B 5120 40960 apache-2.0 none official Qwen3Model — registered 13.99/27.98 GB LMEB rank 30, LME 63.04, LoCoMo 43.76, mean 54.68. RTEB rank 19. LongEmbed full mean 75.57 (rank 6)
microsoft/harrier-oss-v1-0.6b (baseline trial) f9b9dc8d lastMod 2026-03-30 596M 1024 32768 MIT none official (onnx-community conversion = community-only per baseline) Qwen3Model (HF tag confirms qwen3) — registered — this updates the baseline row, which said "community GGUFs unverified"; unlike Nemotron, harrier's bare class IS in the registry, so a self-conversion has a real precedent (Qwen3-Embedding, F2LLM-v2) 0.60/1.19 GB LMEB rank 5 (Borda; 8th by mean, 46th by LongMemEval), LME 72.46, LoCoMo 48.05, mean 58.43 (baseline). LongEmbed: only Passkey scored (76.75%), no full mean
microsoft/harrier-oss-v1-270m — lastMod 2026-03-30 268M 640 32768 MIT none official Gemma3TextModel (HF tag gemma3_text) — registered 0.27/0.54 GB RTEB rank 258. No LMEB row found for 270m specifically
microsoft/harrier-oss-v1-27b — lastMod 2026-03-30 27.0B — 131072 MIT none official Gemma3TextModel — registered 27/54 GB (too large to run comfortably alongside a memory LLM on 64GB) RTEB rank 240. LongEmbed: only Passkey scored (97.5%), no full mean
nvidia/Nemotron-3-Embed-1B-BF16 (baseline trial — the size the baseline itself flags as MPS-deployable) c0c9fea9 created 2026-07-14, lastMod 2026-08-27 1.14B 2048 32768 openmdw-1.1 (notice-preserving + patent-suit termination; not OSI-standard) none official Ministral3Model (bare, base = Ministral-3-3B-Instruct) — NOT registered; only Ministral3ForCausalLM/Mistral3ForConditionalGeneration (causal variants) are (conversion/mistral3.py:14) 1.14/2.28 GB LMEB rank 3 (Borda; 2nd by mean, 3rd by LongMemEval), LME 81.58, LoCoMo 50.77, mean 61.50. RTEB rank 14 (tied w/ Qwen3-Embedding-8B)
nvidia/Nemotron-3-Embed-8B-BF16 (baseline trial) d1f2f257 created 2026-07-14, lastMod 2026-08-28 7.95B 4096 32768 openmdw-1.1 none official; community GGUFs unverified for bidirectional attention (baseline) same Ministral3Model gap 7.95/15.91 GB LMEB rank 1, LongMemEval 85.70 (best found this sweep), LoCoMo 56.51 (best found), mean 64.36. RTEB rank 2/274
nvidia/Nemotron-3-Embed-1B-NVFP4 (baseline trial) f6301282 2026-07-14 / 2026-08-28 705M — — openmdw-1.1 none official same gap 0.70/1.41 GB Not found on LMEB (only BF16 siblings scored) — unknown
Alibaba-NLP/UEmbed-2B / -4B / -9B d02beb2d (2B), 2982eab2 (4B) created 2026-07-29, lastMod 2026-08-18 2.21B/4.54B/8.39B not verified (nested VL config; hidden_size not at top level) not verified cc-by-4.0 none Qwen3_5ForConditionalGeneration nominally registered (conversion/qwen.py:637, text-only class), but repo ships custom src/models/qwen35_embedding.py (trust_remote_code) + separate sparse_weights.pt for a dense+sparse hybrid head the generic class does not model — conversion correctness unverified 2.21/4.43 — 8.39/16.79 GB No LMEB/RTEB/LongEmbed row found. models.md: "conditional," no verified eval
Alibaba-NLP/core-emb-2b / core-emb-8b — created 2026-08-30/31, lastMod 2026-09-04 not verified not verified not verified cc-by-4.0 none not checked unknown No board row; genuinely new (this month). models.md: "watch," paper date only, not a confirmed weight-release announcement
tencent/WeMM-Embedding-2B / -4B / -9B 41a4fafd (4B) created 2026-08-25, lastMod 2026-09-03 2.72B/5.17B/9.41B not verified not verified apache-2.0 (confirmed by reading the LICENSE file itself — the HF metadata tag shows "other" but the license text states Apache-2.0 with third-party-component carve-outs) none same Qwen3_5ForConditionalGeneration caveat as UEmbed 2.72/5.44 — 9.41/18.82 GB No LMEB/RTEB/LongEmbed row — genuinely new (this month)
tencent/R3-embedding-0.6b 9df7e112 created/lastMod 2026-07-08 596M not verified not verified apache-2.0 none Qwen3Model — registered (base_model Qwen3-Embedding-0.6B per README) 0.60/1.19 GB Fine-tuned specifically for agent-skill routing (bi-encoder recall stage of R3-Skill, paired with R3-Rerank-0.6B) — not general passage/memory retrieval; applicability to LongMemEval-style RAG is out-of-domain/unverified. No board row
voyageai/voyage-4-nano — createdAt 2026-01-06 346M not verified not verified apache-2.0 (HF tag; LICENSE.txt file itself not read) none (safetensors + custom modeling_qwen3_bidirectional.py, trust_remote_code) qwen3-tagged bidirectional variant; not individually verified against the registry string 0.35/0.69 GB No RTEB row found for the nano tier specifically (mongodb/voyage-4-large, a separate hosted-only tier, is RTEB(eng,beta) rank 1 overall — not the open-weight model)
jinaai/jina-embeddings-v5-text-small / -omni-small (+nano variants) dd76d535 (text-small) created 2026-01-22/Mar-31, lastMod 2026-04-15/Aug-27 596M–1.63B — 8192–32768 cc-by-nc-4.0 — NON-COMMERCIAL none JinaEmbeddingsV5Model — custom arch, no jina*.py file in conversion/; not registered 0.6–1.6 / 1.2–3.3 GB LMEB rank 28 (text-nano, LME 57.60), rank 50 (text-small, LME 74.70). Double-blocked: NC license + unsupported arch
Linq-AI-Research/Linq-Embed-Mistral — lastMod 2024-06-05 (no change since) 7.11B 4096 32768 cc-by-nc-4.0 — NON-COMMERCIAL none Mistral base, not verified for this bidirectional variant 7.11/14.22 GB LMEB rank 19 by mean but LongMemEval only 38.52 (weak). RTEB rank 105/274. LongEmbed full mean 45.82 (rank 51)
NovaSearch/stella_en_400M_v5 (+1.5B_v5) — lastMod 2025-07-28 (no change since) 435M (/1.5B) 4096 8192 MIT not checked this session not checked 0.44/0.87 GB LMEB rank 12, LME 77.70, LoCoMo 49.03, mean 57.88
Snowflake/snowflake-arctic-embed-l-v2.0 (+m-v2.0) — lastMod 2025-07-28 (l) / 2025-04-24 (m) (no change since) 568M (/305M) 1024 (/768) 8192 apache-2.0 not checked this session not checked 0.57/1.14 GB (l) LMEB rank 15 (l, LME 71.21), rank 13 (m, LME 70.99)
ibm-granite/granite-embedding-311m-multilingual-r2 (+97m, small-english-r2, english-r2) — lastMod 2026-05-18 (r2 pair) / 2026-01-21 (small-en) 312M/97M/48M/149M 768/384/384/768 8192 apache-2.0 not checked this session not checked 0.10–0.31 / 0.19–0.62 GB LMEB ranks 57-91 range (LME 64-69). LongEmbed full mean 71.72 (311m-r2, rank 17) — best-performing tiny official-org model on long-doc after F2LLM-v2

Rerankers found in the sweep (out of scope for the Embedder role, noted for completeness)

Alibaba-NLP/core-reranker-2b / -8b (cc-by-4.0, created 2026-08-30) — not evaluated as an embedder. voyageai/rerank-3 (created 2026-09-23) likewise — added on verification, see Recency below.

Recency, precisely

  • Genuinely new HF repos created in the last 4-6 weeks (Aug 16 – Sep 27, 2026): tencent/WeMM-Embedding-2B/4B/9B (created Aug 25) and Alibaba-NLP/core-emb-2b/8b (created Aug 30). Neither has any LMEB/RTEB/LongEmbed row yet.
  • Added on verification, not text embedders (excluded from the counts above): tencent/EVIE-Preview-4.5B, tencent/EVIE-4.5B, tencent/EVIE-8B (image-document retrievers) and tencent/WeVisDoc-2B/4B (image-document retrievers); voyageai/rerank-3 (a reranker, see the Rerankers section above). None of these change the "no new dense text embedder" conclusion below.
  • New since 2026-07-01 but created before the 6-week window: microsoft/bitnet-embedding-0.6b/270m (Jul 15), tencent/R3-embedding-0.6b (Jul 8), Alibaba-NLP/UEmbed-2B/4B/9B (Jul 29), nvidia/Nemotron-3-Embed-8B/1B-BF16/1B-NVFP4 (Jul 14, already the baseline's own trial rows).
  • Everything else that shows an Aug/Sep lastModified (Nemotron Aug 27-28, jina-v5 Aug 27, UEmbed Aug 18, F2LLM-v2 Sep 3, mixedbread-ai Aug 19, LiquidAI ColBERT Aug 17-18) is a metadata/README revision touch on a model released earlier (F2LLM-v2's actual release was 2026-03-09; the Sep 3 date is not a new drop). Confirmed by comparing each row's createdAt to its lastModified in the raw HF sweep files under raw/*.json.
  • No target-publisher org (Qwen, nvidia, google, microsoft, BAAI, jinaai, Snowflake, ibm-granite, Alibaba-NLP, intfloat, mixedbread-ai, nomic-ai, NovaSearch, Linq-AI-Research, codefuse-ai, tencent, LiquidAI, sentence-transformers, voyageai) released a new dense text embedder in the last 4-6 weeks; the two truly-new repos (WeMM-Embedding, core-emb) are the closest fit and are multimodal/hybrid, unbenchmarked.

llama.cpp architecture registry facts (primary source, this session)

Fetched conversion/*.py from ggml-org/llama.cpp at commit 59657a613ab0fa4ab327d6c790123dff30bfbd67 (catalog b11057) and 7fe450e19305b828c199d602c23a8337aaa1f03b (v0.5.0/b11146); convert_hf_to_gguf.py itself is byte-identical between the two (diff = no output), so "b11057 or later" is one fact set here.

  • conversion/qwen.py:159 — @ModelBase.register("Qwen3ForCausalLM", "Qwen3Model") — covers Qwen3-Embedding, F2LLM-v2 (all sizes), tencent/R3-embedding-0.6b, harrier-oss-v1-0.6b.
  • conversion/qwen.py:637 — @ModelBase.register("Qwen3_5ForConditionalGeneration", "Qwen3_5ForCausalLM") (model_arch=QWEN35, text-only) — covers the base LM class Alibaba-NLP/UEmbed-* and tencent/WeMM-Embedding-* declare, but both ship custom trust_remote_code embedding/sparse heads the generic class does not model; conversion correctness unverified.
  • conversion/gemma.py:182 — @ModelBase.register("Gemma3TextModel") — covers embeddinggemma-300m and harrier-oss-v1-270m/27b.
  • conversion/bitnet.py:11 — @ModelBase.register("BitnetForCausalLM", "BitNetForCausalLM") — consistent with microsoft/bitnet-embedding's native i2_s GGUF.
  • conversion/lfm2.py:68 — @ModelBase.register("Lfm2Model", "Lfm2BidirectionalModel") — covers LiquidAI's embedding/ColBERT lines.
  • conversion/mistral3.py:14 — @ModelBase.register("Mistral3ForConditionalGeneration", "Ministral3ForCausalLM") — does not include bare Ministral3Model, the exact class nvidia/Nemotron-3-Embed-* declares. Primary-source confirmation of the baseline's caution ("community GGUFs are unverified for its bidirectional attention").
  • No conversion/jina*.py file exists — jinaai's v5 custom architecture has no converter.

Ollama library check (this session, HTTP status on ollama.com/library/{name}, guessed tag names)

embeddinggemma=200, granite-embedding=200, qwen3-embedding=200, all-minilm=200; nemotron-embed=404, f2llm=404, bitnet-embedding=404, harrier=404 (a 404 on a guessed name does not rule out a differently-named official tag; not exhaustively searched).

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

S3 input from mac-coordinator-64gb-20260925, 2026-09-27: rerank sweep table (suggested path: inputs/sweep-20260927-mac/rerank.md). Sonnet sweep, independently verified by an Opus pass against primary sources; its "Verification" section lists every correction applied. The published and self-reported numbers here are discovery evidence, not S3 evidence.

Table

Reranker sweep, 2026-09-27 — Mac memory-lane trial candidates

Verification (2026-09-27, independent Opus pass)

  • KaLM-Reranker-V1-Small Mac path reworded: not a "patched llama.cpp fork" but upstream llama.cpp pinned at 277a105 plus 7 patches bundled in the HF repo (llama.cpp/patches/0001-0007); stock llama.cpp does not recognize t5gemma2; only CUDA and CPU builds are documented, no Mac path. Source: verify/rerank/r_kalmgguf_README.md; verify/rerank/m_kalmgguf.json (patch file listing); verify/rerank/g_commit_277a105.json (ggml-org/llama.cpp@277a105).
  • Qwen3-Reranker-0.6B: "verified... b11057-compatible" softened to maintainer-demonstrated on an October 2025 build (b11057 itself untested); footprint corrected to 0.64GB actual (639,153,184 bytes). Source: verify/rerank/g_issue16407.json + g_issue16407_comments.json (Misc. bug: server/rerank output result is wrong with most models include qwen3-Rerank ggml-org/llama.cpp#16407, ggerganov comment 2025-10-03T10:25:20Z); verify/rerank/m_ggml06.json (totalFileSize).
  • BGE-v2-m3 "beats every Qwen3-Reranker size" limited to NDCG@10/MAP (margins +0.0053/+0.0114 over Qwen3-0.6B, no CIs reported); Qwen3-Reranker sizes lead on R@5/R@20 in the same Table 6, so treated as an effective tie; added the one unresolved report of a bge GGUF ranking differently on M1 llama.cpp than on vLLM/CUDA. Source: verify/rerank/a_html_2605.06132v2.html (Table 6, full row incl. R@5/R@20); verify/rerank/g_issue16407_comments.json (heibaidaolx123 comment, 2026-02-02T09:33:48Z).
  • Table 6: GPT-4o-mini was evaluated on 481 of 500 LongMemEval queries (19 skipped on API failures). Source: verify/rerank/a_html_2605.06132v2.html ("GPT-4o-mini evaluated 481 out of 500 queries on LongMemEval").
  • Labeled the "Ranked top 3" list explicitly as a deliberately varied trial lineup (best stock artifact / best memory evidence / lowest-friction ONNX), not the Full table's sort-key order.
  • Added two omitted repos: ggml-org/jina-reranker-v1-turbo-en-GGUF (sha 607d8664, 2024-12-12, apache-2.0, no memory evidence) and Qwen/Qwen3-VL-Reranker-2B/8B (created 2026-01-07, multimodal). Source: verify/rerank/m_ggml_jina_v1t.json; verify/rerank/s_qwen_rerank.json.

Method: HF API (api/models, pipeline_tag=text-ranking, per-org search=rerank, sorted createdAt),
per-model full metadata (api/models/{id}, .sha=40-hex, .lastModified, .tags, .safetensors.total),
raw READMEs via huggingface.co/{id}/raw/main/README.md, arXiv HTML for 2606.22807v3 (KaLM) and
2605.06132v2 (MemReranker), GitHub API for ggml-org/llama.cpp#16407 (state + comments), WebSearch for
MTEB/LightOn context. No ctx_fetch_and_index/ctx_search tool was available in this environment (ToolSearch
found none); curl substituted for it against public read-only HF/GitHub/arXiv endpoints, consistent with
this repo's own pin_method.huggingface convention. No weights downloaded.

Memory footprint: BF16 = params × 2 bytes; Q8_0 ≈ params × 1.06; Q4_K_M ≈ params × 0.6 — "computed" unless
marked "actual" (from a published file size). Mac-suitability sort key, applied in order: (1) runs
unmodified on stock llama.cpp b11057 --reranking, or ships an official MLX build, or an official ONNX
arm64 build; (2) commercial-usable licence; (3) no custom loader / no trust_remote_code; (4)
memory-retrieval-specific evidence (LongMemEval/LMEB), independent weighted over same-team self-reported.

Baseline file: catalogs/foundation/memory-stack-20260925.json — non_repository_decisions rows 1310-1600
(production ai-memory-llm-reranking L1310; MemReranker-4B L1338; KaLM-Reranker-V1-Small L1381;
ettin-reranker-400m-v1 L1425; Qwen3-Reranker-4B L1469; Jina Reranker v3.5 L1511; ms-marco-MiniLM-L6-v2 L1553).

Correction (post first draft): the MemReranker paper's Table 6 has 10 numeric columns
(MAP MRR NDCG@1 NDCG@3 NDCG@10 NDCG R@3 R@5 R@20 F1); the first draft of this file misread column 6
("NDCG", i.e. a different, non-@10 metric) as NDCG@10 for every row except MemReranker-4B. Corrected below.

Production (context, not ranked)

ai-memory LLM listwise reranking = Ollama qwen3.5-9b-64k (qwen3.5:9b-mlx, tag digest 5efc0dd55d4a),
reasoning low, 20s hard timeout. No harness result yet (C4 has not run on the Mac). Footprint: unspecified
exact GB in the catalog (9B nominal, MLX-quantized; not resolved here) — unknown.

Shared evidence table (arXiv 2605.06132 Table 6, LongMemEval n=500, BGE-M3 top-50 candidates, self-reported by MemReranker's authors)

Model MAP NDCG@10
GPT-4o-mini (pointwise LLM judge, hosted; 481/500 queries scored — 19 skipped on API failures, all other rows cover the full 500) 0.5684 0.6265
Qwen3-Reranker-4B 0.6501 0.7185
BGE-v2-m3 0.7069 0.7573
Qwen3-Reranker-8B 0.6966 0.7551
Qwen3-Reranker-0.6B 0.6955 0.7520
Gemini-3-Flash (pointwise LLM judge, hosted) 0.7259 0.7733
MemReranker-0.6B (no published HF repo found — see unknowns) 0.7538 0.7987
MemReranker-4B 0.8043 0.8354

Full table, sorted by Mac suitability (numbered = ranked contenders; unnumbered = out of scope/hosted/rejected)

# Model (HF id) Rev (40-hex) Date (created→lastMod) Params License Mac artifact Custom loader Footprint Strongest evidence
1 Qwen/Qwen3-Reranker-0.6B e61197ed45024b0ed8a2d74b80b4d909f1255473 2025-05-29→2026-04-16 (not in baseline; pre-window) 596M Apache-2.0 ggml-org official GGUF ggml-org/Qwen3-Reranker-0.6B-Q8_0-GGUF rev a02f48bb4f057028298c21fa033da2b30d7742d5 (2025-10-03). Verified, not inferred: llama.cpp maintainer ggerganov posted this exact repo in issue #16407 on 2025-10-03 ("I just uploaded a correct one"), ran llama-server -hf ggml-org/Qwen3-Reranker-0.6B-Q8_0-GGUF ... --rerank and showed well-separated scores (0.999 vs 0.0001-class); issue closed 2025-10-09 No (custom_code:false) 0.64GB actual (639,153,184 bytes) LongMemEval-500: NDCG@10 0.7520/MAP 0.6955 — ties Qwen3-Reranker-8B, beats the 4B, effectively ties BGE-v2-m3 (0.7573/0.7069): BGE leads NDCG@10/MAP by +0.0053/+0.0114 with no CIs reported, but this model leads BGE on R@5/R@20 in the same Table 6 — differs from it mainly in artifact provenance
2 cross-encoder/ettin-reranker-400m-v1 5dca36282a5d85f368d2544002513a29159b4c9e 2026-05-15→2026-05-19 395M Apache-2.0 Official ARM64 ONNX int8 onnx/model_qint8_arm64.onnx, 397.5MB actual No 397.5MB actual; BF16 0.79GB computed MTEB(eng,v2) retrieval-after-rerank 0.6091 vs ms-marco-MiniLM 0.5082; NanoBEIR 0.7193 (DV, no memory-specific number)
3 cross-encoder/ettin-reranker-1b-v1 7d20e9baad17016fdf5549c08f69a2d7ca3e60c3 2026-05-15→2026-05-19 1.03B Apache-2.0 Same family/day as #2; ONNX arm64 sibling assumed, not independently confirmed (siblings fetch returned empty, likely rate-limited) No BF16 2.06GB computed Same family as #2, no size-specific score found
4 cross-encoder/ms-marco-MiniLM-L6-v2 (+ Xenova ONNX) 233902d25c440f23af6f7d6e94d2946bac0bee0a / a09144355adeed5f58c8ed011d209bf8ee5a1fec 2026-08-09 (ce) / 2025-06-30 (Xenova) ~23M (documented) Apache-2.0 (ce copy); Xenova ONNX copy carries no license tag Bundled ONNX by Xenova (community conversion account, not the original publisher), ships inside agentmemory 0.9.29, off by default — flagged community No <100MB 0.5082 (DV, ettin's table) — the floor every candidate above beats
5 BAAI/bge-reranker-v2-m3 953dc6f6f85a1b2dbfca4c34a2796e7dde08d41e 2024-03-15→2024-06-24 568M Apache-2.0 None official found (community ONNX/GGUF conversions exist); it is the model llama.cpp's own server README names as its reference reranking example No BF16 1.14GB computed LongMemEval-500: NDCG@10 0.7573/MAP 0.7069 — leads on NDCG@10 and MAP only (+0.0053/+0.0114 over Qwen3-Reranker-0.6B, no confidence intervals reported); Qwen3-Reranker sizes lead on R@5/R@20 in the same Table 6, so this is effectively a tie rather than beating every Qwen3-Reranker size outright. One unresolved report (llama.cpp issue #16407, comment 2026-02-02) found a bge-reranker-v2-m3 GGUF ranking differently on M1 llama.cpp (Metal) than on vLLM/CUDA — still the reigning "current leader" baseline by reputation
6 IAAR-Shanghai/MemReranker-4B 22890bb4adff85b83b2181ed91e940ffb5a4ec45 2026-04-27→2026-05-12 4.41B Apache-2.0 Community MLX only; no official Mac artifact No (custom_code:false) BF16 8.82GB computed Best in sweep: LongMemEval-500 NDCG@10 0.8354/MAP 0.8043, beats BGE-v2-m3, all Qwen3-Reranker sizes, GPT-4o-mini and Gemini-3-Flash (0.7733) in its own Table 6. Same-team/self-reported (IAAR-Shanghai, GitHub linked to MemTensor/MemOS, correspondence @memtensor.cn)
7 Qwen/Qwen3-Reranker-8B 77d193c791ed757ca307ee72715aa132723da912 2025-05-29→2026-04-16 8.19B Apache-2.0 No official GGUF (community only) No BF16 16.4GB computed LongMemEval-500: NDCG@10 0.7551/MAP 0.6966 — barely ahead of the 0.6B at 13× the size, both behind BGE-v2-m3
8 Qwen/Qwen3-Reranker-4B 22e683669bc0f0bd69640a1354a6d0aebcfeede5 2025-06-03→2026-04-16, unchanged since baseline 4.02B Apache-2.0 No ggml-org/publisher GGUF (only the 0.6B has one); community conversions only, mixed quality No BF16 8.04GB computed MTEB-R 69.76 (DV); LMEB LongMemEval 53.72 (KaLM paper, D); LongMemEval-500 NDCG@10 0.7185/MAP 0.6501 — worst of its own family
9 ibm-granite/granite-embedding-reranker-english-r2 d09d3d6971b689bf9c23839e45a470874d46e13a 2025-08-04→2025-11-06 149.6M Apache-2.0 None official (community ONNX-fp16 fork only, TahaRauf/...) No BF16 0.30GB computed — smallest text-only reranker found BEIR Avg 55.0-55.8, MLDR(en) 44.9-45.8, Miracl(en) 54.2-55.2 (own card; beats ms-marco-MiniLM-L12-v2, bge-reranker-base/large; ties gte-reranker-modernbert-base). No memory-specific number
10 mixedbread-ai/mxbai-rerank-large-v2 ca7e1ee484c37c0ddd8d178a9a5c33cec575c5e6 2025-03-03→2026-04-08 1.54B Apache-2.0 None official found No BF16 3.09GB computed "Current leader" baseline by name recognition; no memory-specific number captured in this sweep
11 zeroentropy/zerank-2-reranker 5eae30d5ee3c6b2df2ef6d723bde45172d761c4c 2025-11-19→2026-07-24 (in-window touch) 4B (documented) Apache-2.0 None official; safetensors only; community GPU-only quant (FP8/compressed-tensors) exists, no Mac build No (card: trust_remote_code no longer required as of May 2026) BF16 ~8GB computed Self-reported NDCG@10 domain avg 0.6714 > Cohere rerank-v3.5 (0.5847) and Gemini-2.5-Flash-listwise (0.5999); "Conversational" domain 0.6140 is a generic proxy, not LongMemEval itself
12 zeroentropy/zerank-1-small-reranker a65fd51c450e9b47fdddab98e31166ecad21af8d 2025-06-30→2026-07-24 (in-window touch) 1.72B Apache-2.0 None official Unverified (not re-checked) BF16 3.44GB computed zerank-1 avg 0.6456 (unclear if this specific "small" size or the full zerank-1-reranker; not disambiguated in fetched text)
13 KaLM-Embedding/KaLM-Reranker-V1-Small (+ -R2) V1: 8c2141c8af7417cd4a58fe7b58f79f1c00ad1d49 (2026-06-02→2026-09-21); R2: 68b74f09ff2ab5df9b9a9a48c812a81aaff8815e (2026-08-14→2026-09-23) see above V1 2.12B total/1B activated (encoder-decoder, T5Gemma2) Apache-2.0 Official GGUF exists but fails the stock-runtime criterion: KaLM-Embedding/KaLM-Reranker-V1-Small-Q8_0-GGUF (91ba8005363b484c48d07eb48b2c35fa5ea0aa99, 2026-07-02) needs upstream llama.cpp pinned at 277a105 plus 7 patches bundled in the Hugging Face repo (llama.cpp/patches/0001-0007) + custom build target llama-kalm-reranker; stock llama.cpp does not recognize the t5gemma2 architecture; only CUDA and CPU builds are documented, no Mac path Yes — ships kalm_cross_encoder.py/kalm_reranker.py/kalm_reranker_utils.py + a vLLM plugin; not flagged by the custom_code tag, confirmed instead by file listing GGUF actual 1.82GB (Q8_0, text params 1.70B); BF16 4.24GB computed LMEB dialogue: V1-Small 87.58 (baseline, DV); V1-Large 89.54 on the LongMemEval sub-task specifically, beating Qwen3-Reranker-8B's 78.93 in the same KaLM paper table (arXiv 2606.22807v3 Appendix G, re-verified in this session — same-team, runs LMEB itself). R2 has no LMEB/LongMemEval number in v3 of the paper — confirmed independently, not just stale from 09-25
14 nvidia/llama-nemotron-rerank-1b-v2 828765652b05bd439c9789d2a6d093db1caa1443 2025-10-16→2026-08-26 (in-window update) 1.24B "other": OpenMDW-1.1 + Llama 3.2 Community Licence; card states "ready for commercial use" but attribution/acceptable-use conditions attach None official (safetensors only) Yes (custom_code:true, trust_remote_code=True in its own usage snippet) BF16 2.47GB computed No BEIR/MTEB number captured in the fetched README excerpt; multilingual (26 languages), 8192-token cross-encoder from Llama-3.2-1B, hosted alternative via NVIDIA NIM

Out of scope / hosted-only / rejected (not numbered — not comparable contenders for this role)

Model Rev (40-hex) Date License Why excluded
jinaai/jina-reranker-v3.5 (+ -mlx, -GGUF) v3.5: e8a93f33f0b22108f8c2364f8484ce3422552fbc; mlx: 3dd4ac901ccdcac85abe3815df0a0aaaf44e4a21; GGUF: 884f7c67aa3ac24edb89064da8c7bfd03f4a90f5 2026-07-14→2026-07-30 CC-BY-NC-4.0 — rejected, non-commercial (baseline, unchanged) Best Mac artifact coverage of the entire sweep (official MLX and official GGUF) but disqualified twice over: non-commercial licence, and newly-confirmed custom_code:true (baseline hadn't flagged this)
lightonai/LightOn-rerank-{PW,LW}-{0.8B,2B,4B} PW-2B: c5333533d5ebea2142490d5eacd42f0c8363c808 (representative) 2026-07-08→2026-07-22 Apache-2.0 Multimodal-first (image-text-to-text, Qwen3.5-VL LoRA): LW-2B nDCG@10 62.66 on ViDoRe V3 (document-image retrieval); "competitive on text-only BEIR" claimed with no exact figure found. No custom loader, would need the VL stack regardless
Alibaba-NLP/core-reranker-8b (+ -2b) 8b: d95d36a34300e34b87cc1793827ac798642c4488 2026-08-30→2026-09-04 CC-BY-4.0 (commercial-OK) N/A for text memory retrieval — a Qwen3-VL-Reranker-based multimodal compositional-image-reasoning reranker (COLA/SugarCrepe++/NegBench/MCMR); no BEIR/LongMemEval text evidence at all. (-2b sibling exists, sha not separately verified)
Contrastive-LM/CLM-v0.1-8B not resolved (excluded before full fetch) 2026-09-21 Apache-2.0 An agent state/action trajectory verifier (DeepSWE 81.6%, Terminal-Bench 2.1 87.6%), tagged text-ranking/reranker but not a document reranker
voyageai/rerank-3 2d2c29517b5c3bbf130c879fda12b1678e598158 2026-09-23 apache-2.0 tag, no weights published Hosted API only (siblings = config/tokenizer, no safetensors/bin)
CohereLabs/rerank-v3.5 d9c89bde0658d73e59e60818bcfcdf8eb388d930 2025-01-02→2026-03-25 none published Hosted API only (3 siblings, no weights); scores 0.5847 avg NDCG@10 in zerank-2's own table — below zerank-1/2 and Gemini-3-Flash
ggml-org/jina-reranker-v1-turbo-en-GGUF 607d8664c787e517e5d6e339d21f680f9002c931 2024-12-12 (pre-window) apache-2.0 Omitted from the original sweep, added on verification. ggml-org official GGUF (jina-bert-v2 arch, 37.6M params, 77MB f16) — otherwise a genuine stock-runtime candidate, but pre-window and no memory-specific evidence found
Qwen/Qwen3-VL-Reranker-2B (+ -8B) 2B: 4bd860ac4f15ad1897a214615cccc700f8f71818; 8B: b212dc8c91a8164aef1ea2de9c1a867611e75c04 2026-01-07 apache-2.0 per the Qwen3-Reranker family, not independently verified for these two VL repos Omitted from the original sweep, added on verification. Multimodal (vision-language) sibling of the Qwen3-Reranker text family — out of scope for text-only memory retrieval

Negative findings (publishers/orgs swept with nothing new since 2026-07-01)

  • BAAI: newest is bge-reranker-v2.5-gemma2-lightweight (2024) and the Matroyshka-ReRanker family (2025-05); no update in the window.
  • mixedbread-ai: mxbai-rerank-{base,large}-v2 last touched 2026-04-08; no v3.
  • Tencent: no rerank-named model under the Tencent HF org.
  • MemTensor: no dedicated reranker product; ships MemOperator-{0.6B,1.7B,4B} and MemReader-4B (2025-07, different function, not verified as rerankers) — but MemReranker-4B (row 6) is affiliated with this same group per its paper's GitHub link and author email domain.
  • answerdotai / jhu-clsp: the ettin family lives under cross-encoder, not these org accounts directly.
  • Qwen: no "Qwen3.5-Reranker" exists — production's own listwise LLM family has no dedicated-reranker sibling on HF.
  • The 4-6 week window (since ~2026-08-16) from named serious text-reranker publishers contains only the KaLM V1→R2 refresh (Aug 14) and Alibaba-NLP's multimodal core-reranker (Aug 30, out of scope) — nothing new from Qwen, BAAI, jina, mixedbread, zeroentropy, or ettin specifically in that narrower window; nvidia's rerank-1b-v2 got a same-repo update Aug 26.

Ranked top 3 for the Mac trial lineup

This lineup is a deliberately varied trial set — best stock artifact (#1), best memory evidence (#2), lowest-friction ONNX (#3) — not the Full table's sort-key order; see that table above for the sort-key ranking.

  1. Qwen/Qwen3-Reranker-0.6B (not in baseline; pre-window release, current-leader pick). Wins: the only model in the entire sweep with an official, unmodified artifact that the llama.cpp maintainer himself demonstrated working — maintainer-demonstrated on an October 2025 build (ggerganov, llama.cpp issue #16407, 2025-10-03); compatibility with the b11057 tag specifically is untested, not verified. Apache-2.0, 0.64GB actual, no custom loader. Loses against production: its LongMemEval-500 number (0.7520 NDCG@10) is competitor-reported, not our own NCE, and it is an effective tie with BAAI/bge-reranker-v2-m3 on that same shared table (BGE ahead +0.0053 NDCG@10/+0.0114 MAP, this model ahead on R@5/R@20) — the two differ mainly in artifact provenance, not a clean evidence-quality win either way. A single small cross-encoder pass is far cheaper per-candidate than a 9B generative listwise call, but Ollama has no rerank endpoint, so either candidate still adds a second server process (llama-server/ONNX runtime) beside Ollama.
  2. IAAR-Shanghai/MemReranker-4B (already a baseline trial; reaffirmed with richer evidence). Wins: the single strongest concrete memory-specific result found anywhere in this sweep — it beats Gemini-3-Flash (a closer analog to an LLM-based reranker than GPT-4o-mini) on the exact same 500-query LongMemEval split, in the paper's own Table 6. Loses: entirely self-reported by its own authors (IAAR-Shanghai/MemTensor); only a community MLX build exists for Mac; baseline's own next-step note stands — it needs "an upstream reranker slot in the chosen memory system; we do not maintain forks" (catalogs/foundation/memory-stack-20260925.json L1349).
  3. cross-encoder/ettin-reranker-400m-v1 (already a baseline trial; reaffirmed as the zero-friction pick). Wins: the only candidate that is both officially packaged for Mac (ARM64 int8 ONNX, 397.5MB, ready today) and needs no patch, no custom loader, no conversion work — lowest integration risk of anything examined. Loses: zero memory-specific or LongMemEval evidence anywhere (this sweep and the baseline agree); its only strength is generic MTEB/NanoBEIR numbers, unproven on the dialogue/temporal-reasoning queries the production reranker actually has to handle.

Close misses, explicitly not top-3: BAAI/bge-reranker-v2-m3 now reads as arguably the strongest all-round candidate on the evidence (ties pick #1 on the shared table — ahead on NDCG@10/MAP, behind on R@5/R@20 — no custom loader, Apache-2.0) but has no confirmed official Mac artifact of any kind, only community conversions — kept at #5 in the ranked table rather than top-3 pending an artifact check. KaLM-Reranker-V1(-R2) has the strongest self-reported raw number (89.54 for V1-Large on LongMemEval specifically) but the worst Mac-artifact story of the sweep — its own GGUF card says stock llama.cpp/Ollama/LM Studio are not supported, and the non-GGUF path needs custom loader code either way; R2 (the checkpoint anyone would actually want) has no memory number at all yet. zerank-2 has a strong generic self-reported benchmark and no trust_remote_code, but no Mac artifact of any kind and no LongMemEval-specific number. No candidate in this sweep satisfies both "official stock-runtime artifact" and "memory-benchmark evidence" at once.

Unknowns

  • MemReranker-0.6B appears in the shared evidence table (NDCG@10 0.7987/MAP 0.7538, second-best in the table) but no matching HF repo was found under the IAAR-Shanghai org listing — the smallest memory-strong reranker in the evidence isn't (yet, or publicly) shipped as weights.
  • ettin-reranker-1b-v1's ONNX arm64 sibling file was not independently confirmed (empty siblings response, likely a rate-limited/short JSON body from the HF API, not re-verified with a second attempt).
  • Whether zerank-2's/zerank-1-small's "Yes/No"-logit classification head converts cleanly through convert_hf_to_gguf.py's Qwen3-Reranker path is unverified — not attempted by anyone found in this sweep.
  • nvidia/llama-nemotron-rerank-1b-v2 and lightonai/LightOn-rerank-* have no BEIR/MTEB/memory number captured from the portions of their cards fetched (README excerpts, not full cards); a further read could surface one.
  • No dedicated ctx_fetch_and_index/ctx_search tool was available in this session (ToolSearch returned only WebSearch and ai-memory's memory_query); HF/GitHub/arXiv reads went through direct curl to public read-only endpoints instead, consistent with this repo's own pin_method.huggingface convention — flagged as a tooling deviation, not a silent one.
  • ai-memory/memory_query was not called: no .ai-memory.toml was found under the working directory to source workspace/project, and none was supplied by the caller; the task's own baseline JSON was used as the source of record instead.
  • MTEB reranking leaderboard (huggingface.co/spaces/mteb/leaderboard) was not scraped beyond one WebSearch pass (JS-heavy Space, no plain-HTML table); no current Fix native worker timeout cleanup on macOS #1 pulled from it directly — current-leader claims here rest on the KaLM/MemReranker papers' own comparison tables instead.
  • LightOn's exact text-only BEIR number (blog says "competitive," no figure) and the precise meaning of its reported "0.6469" for the 4B variant (scale/benchmark unstated in the fetched snippet) were not resolved.
  • Alibaba-NLP/core-reranker-2b's full 40-hex revision was not re-fetched after the first truncated read (excluded from the role anyway; low priority to close).

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

S3 input from mac-coordinator-64gb-20260925, 2026-09-27: memllm sweep table (suggested path: inputs/sweep-20260927-mac/memllm.md). Sonnet sweep, independently verified by an Opus pass against primary sources; its "Verification" section lists every correction applied. The published and self-reported numbers here are discovery evidence, not S3 evidence.

Table

Local memory-LLM sweep, 2026-09-27 (since 2026-07-01, emphasis last 4-6 weeks)

Verification (2026-09-27, independent Opus pass)

  • Gemma-4 license: removed the "custom Gemma terms" caveat everywhere it appeared (rows 2 and 6, Unknowns, ranking (a) items 2/3) - license_link redirects to a plain "Apache License 2.0" page, and the card body itself says "License: Apache 2.0". Source: verify/memllm/raw/gemma_4_license.html (canonical ai.google.dev/gemma/apache_2), verify/memllm/raw/README_google_gemma-4-26B-A4B-it.md, verify/memllm/raw/api_google_gemma-4-26B-A4B-it.json (cardData.license_link).
  • LFM2.5-8B-A1B does have an Ollama tag (row 1, ranking (a) item 1): lfm2.5:8b and lfm2.5:8b-a1b-q4_K_M both resolve; added the LFM Open License v1.0 terms (US$10M revenue threshold incl. >=50%-controlled affiliates, free non-commercial/local use, automatic termination on breach). Source: verify/memllm/raw/ollama_lfm2.5_tags.html, ollama_lfm2.5_8b_manifest.json; verify/memllm/raw/LICENSE_LiquidAI_LFM2.5-8B-A1B (Sections 1, 5, 11).
  • granite-4.2-8b BFCL v4 corrected to 52.39 at the pinned revision f8de16cdcdbc6c779ca517604e050d82cc119e44 (row 3, Sources, ranking (a) item 3) - 50.29 was a superseded 2026-08-18 card. Source: verify/memllm/raw/commits_granite8b.json (commit dated 2026-08-18T19:55:56Z superseded by 2026-09-04T21:02:25Z), verify/memllm/raw/src_eesel.html ("8B is close enough (52.39 BFCL...)").
  • Granite speed (row 4, Unknowns, ranking (b) item 3): granite-4.2-30b is 76.6 tok/s on Artificial Analysis's hosted API (not Mac), AA's own label "slower than average" - not "notably slow, 28 tok/s"; granite-4.2-8b is 84.9 tok/s (Unknowns). "Notably slow" is AA's wording for Qwen3.8 (xhigh) at 45.5 tok/s. Source: verify/memllm/raw/src_aa_gr30.html, src_aa_gr8.html, src_aa_qwen38.html (each page's own stat tile).
  • Dropped "the only hard instruction-following number" (ranking (a) item 1): Liquid's own comparison card gives Gemma-4-26B-A4B-it IFEval 91.40/IFBench 47.25; AA independently records Gemma IFBench 72.4/45.4 and LFM2.5-8B-A1B approx 55.7. Source: verify/memllm/raw/README_LiquidAI_LFM2.5-8B-A1B.md (Knowledge-and-instruction-following table); verify/memllm/raw/src_aa_gemma26.html (embedded chart JSON, ifbench field).
  • Credited "Qwen3.8 beats all medium models / ties DeepSeek-V4-Flash" (row 5, ranking (b) item 1) to a Hacker News comment on an older Intelligence Index scale, not to Artificial Analysis. Source: verify/memllm/raw/src_hn.html (https://news.ycombinator.com/item?id=49334544).
  • LFM2.5-8B-A1B (public 2026-05-28, row 1) and Gemma-4-26B-A4B-it (public by 2026-04-01, row 2) flagged as released before this sweep's since-2026-07-01 window; only their later touches (Aug/Jul) fall inside it. Source: verify/memllm/raw/api_LiquidAI_LFM2.5-8B-A1B.json, api_google_gemma-4-26B-A4B-it.json (createdAt).
  • Replaced the "AA-LCR unrecoverable" unknown with the recovered AA-LCR v1.1 scores, reproducing the baseline's Qwen3.5-9B 46.0/70.0, Qwen3.6-35B-A3B 64.3/71.7 and Nemotron 60.3 exactly as a cross-check (Unknowns; rows 2, 4, 5, 6; ranking (b) fully rewritten on this metric). Source: verify/memllm/aa_lcr_parsed.tsv.
  • Rewrote ranking (b) in full on AA-LCR (the table's own deciding metric for that role): new order Qwen3.8-27B > Gemma-4-26B-A4B-it > granite-4.2-30b; only Qwen3.8-27B beats both the control and the trial. Re-decided ranking (a)'s Publish evidence-led convergence practice and retrieval comparisons #2/Add searchable ecosystem manifest and token-efficiency guide #3 now that Gemma's license caveat is gone: Gemma-4-26B-A4B-it moves to Publish evidence-led convergence practice and retrieval comparisons #2 on this role's own stated discriminator (throughput: 3.8B active vs granite-8b's 8.8B fully dense), granite-4.2-8b to Add searchable ecosystem manifest and token-efficiency guide #3; reasons stated inline. Source: verify/memllm/aa_lcr_parsed.tsv plus the row-level sources above.

Method: Hugging Face API (/api/models?author=...&sort=createdAt, then per-repo /api/models/{id} for
sha/lastModified/license/gated/safetensors.total, then resolve/main/config.json) for architecture;
llama.cpp b11057 = commit 59657a613ab0fa4ab327d6c790123dff30bfbd67 (GitHub tree + raw conversion/*.py at that
commit, matches the catalog pin exactly); Ollama support = HEAD registry.ollama.ai/v2/library/<name>/manifests/<tag>
(200/404, read-only, no pull); WebSearch for AA/IFEval/BFCL; primary LICENSE/README/commit-history reads for the
flagged claims below. All revisions are 40-hex shas observed now via the HF API on 2026-09-27 (see Revisions).
Q4/Q8 sizes are exact GGUF blob bytes from ?blobs=true unless marked otherwise.

Correction after review: config.json shows the production control Qwen3.5-9B is itself
Qwen3_5ForConditionalGeneration / pipeline_tag image-text-to-text (same wrapper class as Qwen3.8-27B). "VL-wrapped"
is dropped as a loss reason for Qwen3.8-27B and softened for Gemma-4-26B-A4B - not a differentiator from the running control.

Baseline (catalogs/foundation/memory-stack-20260925.json, non_repository_decisions, role "Memory LLM")

Model Repo Date Params License Mac artifact Evidence
Qwen3.5-9B (retain, control) Qwen/Qwen3.5-9B 2026-03-02 9B, qwen3_5, image-text-to-text tag Apache-2.0 Ollama qwen3.5:9b-mlx 203e30078279 / prod tag qwen3.5-9b-64k 5efc0dd55d4a D: AA-LCR 46.0 / 70.0(reason). DV LongBench-v2 55.2
Qwen3.6-35B-A3B (trial) Qwen/Qwen3.6-35B-A3B 2026-04-24 35B/A3B MoE Apache-2.0 Ollama qwen3.6:35b-a3b 096fdbd02fe6(GGUF 22.6GB)/-nvfp4 e92a3e94bbca(MLX 23.6GB) D: AA-LCR 64.3/71.7 (+18.3 vs control). No LongMemEval
LFM2.5-2.6B (trial) LiquidAI/LFM2.5-2.6B-GGUF 2026-09-22 2.6B dense LFM1.0 (custom, flag) no Ollama tag (404) DV IFStruct 85.49; 220 tok/s M5 Max CPU
Nemotron-3.5-Lightning-30B-A3B (defer) nvidia/...-BF16 2026-08-24 30B/A3B MoE OpenMDW-1.1 Ollama nemotron-3.5-lightning:30b-a3b e7a64ff15fb1 D: AA-LCR 60.3(reason only). DV 52.0. In-window artifacts: Base-BF16 (08-05), NVFP4/DSpark/DFlash (08-04/05) - Blackwell NVFP4 only, no new Mac evidence, decision unchanged
MemReader-4B-thinking (reject) IAAR-Shanghai/MemReader-4B-thinking 2026-04-08 4B Apache-2.0 - DV LongMemEval 83.0% (4-tool pipeline, not standalone)

New finds since 2026-07-01, in 3-35B class, text-capable (sorted by suitability)

# Model Date (created/lastMod) Params License Mac artifact llama.cpp b11057 Q4/Q8 (exact bytes) Ctx Key evidence
1 LFM2.5-8B-A1B 2026-05-28/08-24 (public 2026-05-28, before this sweep's since-2026-07-01 window; only the Aug touches fall inside it) (DSpark 08-10, DSpark-GGUF 08-19) 8.47B/~1.5B active MoE (Lfm2MoeForCausalLM, 32 experts×top4) LFM Open License v1.0: Commercial Use is unlicensed for a Legal Entity (including affiliates it controls, or that control it, at >=50% ownership) at or above US$10,000,000 annual revenue; non-commercial or local use is free regardless of revenue; any breach terminates the licence automatically official GGUF+MLX(4/5/6/8bit,bf16)+ONNX+unsloth mirror; Ollama tag confirmed: lfm2.5:8b and lfm2.5:8b-a1b-q4_K_M both resolve (200) — corrects the earlier 404 finding Lfm2MoeForCausalLM registered, conversion/lfm2.py:97 5.16GB/9.01GB 128k D/DV: IFEval 91.84 (+12.4 over LFM2-8B-A1B's 79.44); vendor claims IFEval parity w/ Gemma-4-26B-A4B. No LongMemEval. AA-LCR: 0 recorded on Artificial Analysis (a literal zero, not a missing value — genuineness unverifiable)
2 Gemma-4-26B-A4B-it 2026-03-11/07-20 (public by 2026-04-01, before this sweep's since-2026-07-01 window; only the 07-20 touch falls inside it) 25.2B/3.8B active MoE (Gemma4ForConditionalGeneration, 128 experts×top8) Apache-2.0, plain: the README's license_link (ai.google.dev/gemma/docs/gemma_4_license) redirects to ai.google.dev/gemma/apache_2 ("Apache License 2.0"), and the card's own body text reads "License: Apache 2.0" - corrects the earlier custom-Gemma-terms caveat; no restriction beyond Apache-2.0 official Google QAT Q4_0 GGUF (14.44GB text+1.19GB mmproj); ggml-org/gemma-4-26B-A4B-it-GGUF; lmstudio-community GGUF+MLX-4bit; Ollama gemma4:26b (200) Gemma4ForConditionalGeneration registered, conversion/gemma.py:631 Q4(official QAT Q4_0)=14.44GB+1.19GB mmproj; Q8(ggml-org, exact)=26.86GB+0.81GB mmproj 260k D: AA Intelligence Index v4.3.2 (reasoning, est.) ~17 - different benchmark version than baseline's plain AA-LCR. AA-LCR v1.1 recovered on verification: 65.7 (reasoning) / 42.3 (non-reasoning) - see Unknowns. D-third-party (dev.to/aurigait): "a July 2026 refresh" reportedly cut JSON/parameter errors (Tau2 Telecom +10.1%, TB2 +4.5% on sibling 31B) - checked the repo's own commit log: the 07-20 commit is "Add response_template to tokenizer_config.json," 07-15 is a chat-template fix ("null handling, reasoning preservation, turn-tag balance, input validation"); these are template/config commits consistent with, but not confirmed as, the specific change the third-party figures describe. Sibling Gemma-4-31B: 0.798 Value-Accuracy, independent Structured-Output-Benchmark (arXiv 2604.25359). Liquid's own comparison card also lists this model's IFEval 91.40 / IFBench 47.25; Artificial Analysis independently records IFBench 72.4 (reasoning) / 45.4 (non-reasoning)
3 granite-4.2-8b 2026-08-07/09-04 8.79B dense (GraniteForCausalLM) Apache-2.0 official ibm-granite GGUF+MLX(4/6/8bit); lmstudio-community GGUF+MLX; Ollama granite4.2:8b (200) GraniteForCausalLM registered, conversion/granite.py:17 (long-supported) 5.35GB/9.35GB 128k DV: BFCL v4 tool-calling 52.39 at the pinned revision (f8de16cdcdbc6c779ca517604e050d82cc119e44, 2026-09-04) - corrects 50.29, which was a superseded 2026-08-18 card; eesel.ai's independent review reports 52.39 too; native "reason-before-call" tool design; thinking on/off switch
4 granite-4.2-30b 2026-08-07/09-04 29.28B dense Apache-2.0 same coverage; Ollama granite4.2:30b (200) registered, same class 17.72GB/31.11GB 128k DV: BFCL v4 61.39 (best of new finds). D (Artificial Analysis hosted API, not a Mac measurement): 76.6 tok/s, AA's own tier label is "slower than average" - corrects an earlier "notably slow, 28 tok/s" misreading; AA reserves "notably slow" for Qwen3.8 (xhigh) at 45.5 tok/s, not this model. AA-LCR v1.1 recovered on verification: 49.0 - see Unknowns. Latency risk vs the 20s gate remains (fully dense, no MoE discount), just less severe than 28 tok/s implied
5 Qwen3.8-27B 2026-08-05/08-14 27.78B dense, Qwen3_5ForConditionalGeneration (same wrapper class as the production control) Apache-2.0 unsloth+lmstudio-community GGUF(Q4_K_M-UD 16.46GB/Q8_0 29.05GB, exact)+MLX(4/5/6/8bit); Ollama qwen3.8:27b (200) Qwen3_5ForConditionalGeneration/Qwen3_5MoeForCausalLM registered, conversion/qwen.py:637,643 16.46GB/29.05GB (exact) 260k D: AA Intelligence Index (xhigh) 34 - composite index, not plain AA-LCR. AA-LCR v1.1 recovered on verification: 82.0 (xhigh) / 79.7 (medium) / 77.3 (low reasoning), 69.3 (non-reasoning) - the only new find that beats both the control (46.0/70.0) and the trial (64.3/71.7) on AA-LCR, on both axes; see Unknowns. "Beats all medium 40-150B models, ties DeepSeek-V4-Flash-0731(304B)" is a Hacker News comment's reading of an older Intelligence Index scale (news.ycombinator.com/item?id=49334544), not an Artificial Analysis claim
6 Gemma-4-12B-it 2026-05-23/07-20 11.96B dense, omni (Gemma4UnifiedForConditionalGeneration) Apache-2.0, plain (same correction as row 2 - the license_link caveat is resolved, not a restriction) official Google QAT Q4_0 GGUF (6.98GB+0.175GB mmproj); ggml-org/gemma-4-12B-it-GGUF; lmstudio-community GGUF+MLX-4bit+QAT-GGUF; Ollama gemma4:12b (200) Gemma4UnifiedForConditionalGeneration registered, conversion/gemma.py:812,933 Q4(official QAT)=6.98GB+0.175GB; Q8(ggml-org,exact)=12.67GB+0.159GB 260k D: AA Intelligence Index non-reasoning 20 / reasoning 14 (v4.3.2, not plain AA-LCR). AA has a direct release-comparison page vs Qwen3.5-9B (Sources). Closest in-class size analogue to the control. AA-LCR v1.1 recovered on verification: 63.7 (reasoning) / 35.0 (non-reasoning) - see Unknowns
- granite-swash-3b-a600m (excluded: base-only, short ctx) card date 2026-07-07 3.02B/0.598B active MoE (GraniteMoeSWAForCausalLM, 48×top4, SWA+sinks) Apache-2.0 none found GraniteMoeSWAForCausalLM registered, conversion/granite.py:124 n/a 8192 only README states verbatim: "early exploration...small-scale preview for upcoming Granite series," explicitly a base model, not instruct - JSON/tool reliability unproven
- ContextPilot-14B/8B/E4B (excluded: license) 2026-08-27/08-31 14.77B dense (Qwen3ForCausalLM) LICENSE file read directly (primary source, not a paraphrase): body is Apache License 2.0 text with an added "Section 0": "ContextPilot-14B is made available solely for the purpose of scientific research and development. You shall not use it for any other purpose." Hard research-only restriction, confirmed, blocks adoption community-only: bartowski/mradermacher GGUF (not ggml-org/unsloth/lmstudio-community); no Ollama tag (404) Qwen3ForCausalLM long-supported - 40960 native DV, likely self-reported (arXiv 2608.28476, Tencent authors): 72.20 avg on 4 long-context benchmarks. On-theme (proactive context mgmt, structured memory, retrieval, "soft context offloading") but license-blocked
- Nemotron-Labs-Audex-30B-A3B (excluded: wrong modality) created 2026-07-06 ~30B (shards) other - - - - File tree = whisper/audiogen/VAE/speech-decoder shards - audio S2S system, not a text memory LLM despite MoE-like name
- granite-4.2-3b (in-class, not ranked - too small for consolidation, notable for extraction speed) 2026-08-07/09-02 3.66B dense Apache-2.0 official+lmstudio-community GGUF/MLX; Ollama tag not probed registered 2.24GB/3.89GB 128k Same family as rows 3-4, no separate benchmark found

Oversized "Flash"-branded releases this window (excluded on total-footprint alone; 64GB unified memory)

Model Date Total params (safetensors, exact) License tag
Qwen/Qwen3.8-Flash-Next 2026-08-24/27 180.0B other
zai-org/GLM-5.3-Flash 2026-08-25/09-07 321.3B mit (tag)
zai-org/GLM-5.3 2026-08-25/09-04 753.3B other
deepseek-ai/DeepSeek-V4-Flash-0731 2026-07-31/08-01 304.2B mit (tag)
XiaomiMiMo/MiMo-V2.6-Flash-RL 2026-09-21/22 310.8B mit (tag)

Memory-specialised releases: dedicated sweep

Searched: (1) IAAR-Shanghai (MemReader-4B-thinking's org) full listing; (2) search=MemReader; (3) search=memory&pipeline_tag=text-generation,
filtered createdAt>=2026-07-01, limit 100. Results: IAAR-Shanghai's own follow-up to MemReader is MemPrivacy-4B/1.7B-RL/SFT (2026-05-08/09,
privacy-focused, not extraction/consolidation) and MemReranker-4B (2026-04-27, a text-classification cross-encoder, not a generative LLM) -
both predate the window; IAAR-Shanghai's only in-window releases are Metis-4B/9B/27B (VL, 2026-07-16, no memory framing). The generic memory-keyword
search surfaced ~35 in-window hits, all small independent/research repos with no vendor backing or benchmark evidence meeting the "serious" bar:
domain-siloed SFT variants (Jiarui-Wang/MemSFT-Qwen3-{Bio,Law,OpenSWI}-Memory-{1.7B..8B}, 2026-08-01), architecture probes
(Rubin-Wei/MemoryDecoder-{OLMo,Qwen3,Pythia}-*, 2026-07-23), and single-author projects (Jepoxy/LFM2.5-350M-Memory-Extractor,
vtava/Laya-MemoryFusion-*, OLAResearchX/memoryathena-*). None outranks ContextPilot (above, license-excluded) or the generalist rows 1-6 on
any documented metric. No serious vendor-backed memory-specialised LLM release found in-window beyond ContextPilot.

Orgs checked with nothing new/relevant since 2026-07-01

  • openai: gpt-oss-20b/120b unchanged since 2025-08-26 (lastModified); no successor in any window.
  • moonshotai: only Kimi-K3 touched (lastMod 2026-09-02, createdAt 2026-06-13, pre-window, VL, oversized).
  • meta-llama: nothing since Llama-4/Llama-Guard-4 (2025-04/05); stalest org checked.
  • baidu: nothing since ERNIE-4.5 updates (2025-11-26); latest any touch is an OCR model (2026-07-29).
  • stepfun-ai (org is stepfun-ai, not stepfun): nothing since Step-3.7-Flash (2026-05-23, VL, pre-window).
  • allenai: only climate/science models in window (ACE2S, SamudrACE) - no LLM.
  • google, tencent, nvidia, zai-org, deepseek-ai, XiaomiMiMo: covered above; remainder of their Aug-Sep 2026 output is OCR/embedding/rerank/UI-agent/TTS/robotics/vision, not general memory LLMs.

Revisions (40-hex, HF API .sha, resolved 2026-09-27)

Baseline: Qwen/Qwen3.5-9B c202236235762e1c871ad0ccb60c8ee5ba337b9a · Qwen/Qwen3.6-35B-A3B 995ad96eacd98c81ed38be0c5b274b04031597b0 ·
LiquidAI/LFM2.5-2.6B-GGUF e7caca5d835a3901a8e0d63e94009429bafafdfc (base LiquidAI/LFM2.5-2.6B 654f9463ce32b05d0429d76fe1f580b27d4c1ac0) ·
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 a9904d24bcc1d289a1950fa9d2b978c47cf903b9 · IAAR-Shanghai/MemReader-4B-thinking 3fcb57e5653ec5d733eea4df61f4b339d4ac0e92

New finds: LiquidAI/LFM2.5-8B-A1B 5dd22602c2e9f6a097b1de4c4efe0658b605015c · google/gemma-4-26B-A4B-it 4d7ae4984b7db7de8f8457170b3f1a419ee76d52 ·
google/gemma-4-12B-it 707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7 · ibm-granite/granite-4.2-8b f8de16cdcdbc6c779ca517604e050d82cc119e44 ·
ibm-granite/granite-4.2-30b 9e668ce1c538387ef24d3644e9b0606647762636 · ibm-granite/granite-4.2-3b e459acceac81e5fe67c07d9cfc72329a332e7eb1 ·
Qwen/Qwen3.8-27B 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 · ibm-granite/granite-swash-3b-a600m 568a94fcfeff42a8cab9ee2ffea89fde12071282 ·
tencent/ContextPilot-14B 8eafd8356c5c5e713e92e623c8f1d346de6d0f95 · nvidia/Nemotron-Labs-Audex-30B-A3B 4e7e342045736382ddf3e2952c313847a08642b8

Oversized: Qwen/Qwen3.8-Flash-Next de4b8e4d43b917e7706784d8bb445c9af86a3540 · zai-org/GLM-5.3-Flash eb9eb208eb0d988989d07a6a12d0fdeb5f52574a ·
zai-org/GLM-5.3 aca966e4e02791568aa6a4ced368624b3d897f42 · deepseek-ai/DeepSeek-V4-Flash-0731 7872f01b1d1fe23eabc4c98b48bffcef5a386062 ·
XiaomiMiMo/MiMo-V2.6-Flash-RL 5711b268169967567844e1e560e8a3966da959b1

llama.cpp b11057 pin: https://github.com/ggml-org/llama.cpp/tree/59657a613ab0fa4ab327d6c790123dff30bfbd67 (tag b11057, matches catalog).

Sources (by row)

Unknowns / evidence gaps

  • AA-LCR v1.1 recovered on verification (the client-side-rendered score the first pass could not extract via curl) - recovered directly from the Artificial Analysis page data, not re-derived: Qwen3.5-9B (control) 46.0/70.0(reason), Qwen3.6-35B-A3B (trial) 64.3/71.7, Nemotron-3.5-Lightning 60.3(reason only) all reproduce the baseline/table exactly, cross-validating the recovery method. New finds: Qwen3.8-27B 82.0 (xhigh) / 79.7 (medium) / 77.3 (low reasoning), 69.3 (non-reasoning); Gemma-4-26B-A4B-it 65.7 (reasoning) / 42.3 (non-reasoning); Gemma-4-12B-it 63.7 (reasoning) / 35.0 (non-reasoning); granite-4.2-30b 49.0; granite-4.2-8b 45.0; LFM2.5-8B-A1B 0 (recorded, a literal zero - genuineness unverifiable). Only Qwen3.8-27B beats both the control and the trial on this metric; Gemma-4-26B-A4B-it and granite-4.2-30b both trail the control itself now that their scores are known.
  • No LongMemEval / consolidation-specific / listwise-reranking benchmark found for any new candidate (same gap the
    baseline records for Qwen3.6-35B-A3B itself).
  • LFM2.5-8B-A1B active params: name implies ~1B, Liquid's blog says "~1.5B" - both reported, not reconciled.
  • p95 latency at 4-8K tokens (ai-memory's 20s/5%-fallback gate) was not measured for any candidate; qualitative AA
    speed figures (Artificial Analysis's hosted API, not a Mac measurement) are granite-4.2-30b 76.6 tok/s, AA's own
    label "slower than average" (not "notably slow" as an earlier draft had it - AA reserves "notably slow" for
    Qwen3.8 (xhigh) at 45.5 tok/s), and granite-4.2-8b 84.9 tok/s.

Ranked top 3

(a) Hindsight/MemPalace extraction (runs on every retain; throughput + JSON reliability dominate)

  1. LFM2.5-8B-A1B, conditional on clearing the LFM1.0 license gate (free for non-commercial/local use; commercial
    use unlicensed at/above US$10M annual revenue including >=50%-controlled affiliates; any breach terminates it).
    ~1.5B active beats Qwen3.5-9B's 9B-all-active on per-token cost, and IFEval 91.84 is a strong instruction-following
    number (neither baseline model has one on record) - though not the only one surfaced: Gemma-4-26B-A4B-it also
    carries IFEval 91.40/IFBench 47.25 on Liquid's own comparison card, and AA-LCR is now known for this model too
    (0, recorded - genuineness unverifiable). Vs Qwen3.6-35B-A3B: fewer active params still (1.5B vs 3B), a plausible
    speed win, but zero memory/consolidation evidence (the trial 35B-A3B at least has the +18.3 AA-LCR proxy); Ollama
    now carries a tag (lfm2.5:8b, confirmed 200 - corrects the earlier 404 finding).
  2. Gemma-4-26B-A4B-it - the deciding factor for this role is throughput, and it wins that outright: 3.8B active
    vs granite-4.2-8b's 8.8B fully dense (no speed win over the control at all). Now unambiguous Apache-2.0 (the
    license_link caveat is resolved, no longer a differentiator against granite). Carries IFEval 91.40/IFBench 47.25
    on Liquid's own card (AA independently records IFBench 72.4/45.4) - a closer proxy to JSON-instruction reliability
    than granite's tool-calling BFCL. AA-LCR: this model's 65.7 reasoning / 42.3 non-reasoning vs granite-4.2-8b's
    newly-recovered single-mode 45.0 - Gemma ahead on the reasoning comparison, granite slightly ahead against Gemma's
    non-reasoning figure; not a clean win for either, so throughput remains the deciding factor above. Caveat: fits
    this role only at Q4 (14.44GB+1.19GB mmproj) - the Q8 build (26.86GB+)
    exceeds the 35B-A3B trial's own footprint.
  3. granite-4.2-8b - same size class as Qwen3.5-9B (8.8B vs 9B dense, no speed win over the control), but new
    (Sept 2026), unambiguous Apache-2.0, and ships fresh native tool-calling evidence (BFCL v4 52.39 at the pinned
    revision) the baseline lacks - the only BFCL number of the two. Vs Qwen3.6-35B-A3B: loses the active-compute race
    (8B active > 3B active) - wins only if Mac memory headroom, not latency, is binding (5.35GB vs 22.6GB Q4 on disk).
    Demoted behind Gemma-4-26B-A4B-it now that Gemma's license caveat is gone and this role's own stated discriminator
    (throughput) favors Gemma's MoE by more than 2x.

(b) ai-memory consolidation/reranking (bounded 20s / <=5% fallback at 4-8K; quality-bound)

Rewritten on AA-LCR, which this table calls the deciding metric for this role, now that the v1.1 scores are
recovered (see Unknowns) instead of "not recoverable." Non-reasoning: Qwen3.8-27B 69.3 > trial Qwen3.6-35B-A3B 64.3 >
control Qwen3.5-9B 46.0 > Gemma-4-26B-A4B-it 42.3. Reasoning: Qwen3.8-27B 82.0 > trial 71.7 > control 70.0 >
Gemma-4-26B-A4B-it 65.7 > granite-4.2-30b 49.0. Only Qwen3.8-27B beats both the control and the trial; Gemma and
granite-30b both now measurably trail the control itself, which was hidden while their scores were unknown.

  1. Qwen3.8-27B - the only candidate that clears the deciding metric: AA-LCR 82.0 (xhigh) / 79.7 (medium) / 77.3
    (low reasoning), 69.3 (non-reasoning), beating both the control (46.0/70.0) and the trial (64.3/71.7) on every
    reasoning tier and non-reasoning. Same qwen3_5 lineage and same ForConditionalGeneration wrapper as the
    production control (prompt/tokenizer compatibility, not a new liability - this was confirmed, not assumed), 260k
    context. "Beats all medium 40-150B models" is a Hacker News reading of an older Intelligence Index scale, not an
    AA claim - dropped as evidence here. Open risk: fully dense (27.78B active vs the trial's 3B active), and p95
    latency against the 20s gate is unmeasured for this model on this Mac - the one thing standing between this pick
    and displacing the trial.
  2. Gemma-4-26B-A4B-it - longer native context (260k) than the control, MoE 3.8B active vs the control's
    9B-all-active, and the widest Mac tooling of any new find (official QAT GGUF + ggml-org + lmstudio-community +
    Ollama tag, vs the control's single Ollama MLX tag); now unambiguous Apache-2.0. On the deciding metric it loses
    to both the control and the trial (65.7/42.3 vs 70.0/46.0 and 71.7/64.3) - a materially weaker position than "no
    AA-LCR number was recoverable" implied - but it still clears granite-4.2-30b by a wide reasoning margin (65.7 vs
    49.0). A third-party JSON-reliability-refresh claim exists for this window but is not confirmed against the
    primary commit log
    (row 2).
  3. granite-4.2-30b - best tool-calling score of any new find (BFCL v4 61.39), 128k context, unambiguous
    Apache-2.0. On AA-LCR (49.0) it is the weakest of the three role candidates, trailing the control on reasoning by
    21 points. Vs the 35B-A3B trial: similar total disk (30B vs 35B) but fully dense (30B active vs 3B active);
    AA's hosted-API speed figure is 76.6 tok/s labeled "slower than average" (not "notably slow" - that label is
    AA's wording for Qwen3.8 xhigh at 45.5 tok/s) - a real but smaller latency risk against the 20s p95 gate than
    the earlier 28 tok/s figure implied, and one the MoE trial does not carry at all.

Operating point: at Q4, every ranked row (1-6) fits well under the Qwen3.6-35B-A3B trial's 22.6GB GGUF footprint
(largest is granite-4.2-30b at 17.72GB); at Q8, rows 2 (26.86GB+), 4 (31.11GB) and 5 (29.05GB) exceed it - this is the
memory-headroom argument behind picks (a)#2 and (b)#2, and it only holds at Q4/QAT-Q4_0, not at Q8.

Evidence classes follow the catalog's own D (discovery/independent) / DV (discovery/vendor) convention; every score
above is one or the other, never NCE (no native Mac harness run was performed in this sweep).

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

S3 input from mac-coordinator-64gb-20260925, 2026-09-27: gen sweep table (suggested path: inputs/sweep-20260927-mac/gen.md). Sonnet sweep, independently verified by an Opus pass against primary sources; its "Verification" section lists every correction applied. The published and self-reported numbers here are discovery evidence, not S3 evidence.

Table

Local-generation model sweep, 2026-08-01 to 2026-09-27 (revised after review)

Verification (2026-09-27, independent Opus pass)

  • Nemotron-3.5-Lightning-30B-A3B KV at 32K recomputed from the real layer split (6 attention / 23 mamba / 23 moe of 52 layers, 2 kv heads x head_dim 128) instead of an all-52-layers-as-attention upper bound: 0.20GB, not <=1.74GB; total revised to approx 27.0GB (25.27 + 0.20 + 1.49), below granite-4.2-30b's 27.8GB, so the two rows are swapped into ascending-fit order and the "KV split not resolved" unknown is removed. Source: verify/gen/cfg_nemotron_bf16.json (layers_block_type).
  • Corrected the two HF slugs that 401 under their short form to the slugs that actually resolve: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 and prism-ml/Ternary-Bonsai-2-27B-gguf. Source: verify/gen/api_nvidia_NVIDIA-Nemotron-3.5-Lightning-30B-A3B.json and api_prism-ml_Ternary-Bonsai-2-27B.json (both {"error":"Invalid username or password."}) vs. api_nvidia_NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16.json / api_prism-ml_Ternary-Bonsai-2-27B-gguf.json (both resolve).
  • Bonsai 2 27B: stock llama.cpp rejects the shipped PQ2_0/PTQ1_0 files outright (unknown types); "garbage output" is specific to the separate, generic Q2_0 type it silently loads with no Hadamard runtime — these are not the same claim. Added the card's own M5 Pro throughput (approx 28 tok/s), the Mac target chip, ahead of the M5 Max aside. Source: verify/gen/readme_bonsai2_27b_gguf.md lines 139-141 (rejection wording), 191 and 206 (M5 Pro tok/s).
  • Qwen3.8-27B's "smallest 32K footprint" claim qualified to "among rows with a worked-out KV" — Muse-Glimmer-30B's KV was never resolved in this sweep, so it is not a counted comparison (see this table's own Unknowns).
  • Newest llama.cpp tag is b11214 (2026-09-27T12:45:21Z); reconfirmed Xing4_0ForCausalLM absent at both b11146 and b11214 via a fresh architecture-registry scan (previously only checked at b11057/b11214). Source: verify/gen/archs_b11146.txt, archs_b11214.txt, scan2_b11146.txt, scan2_b11214.txt (0 files contain xing4 at either tag).
  • Hosted models: confirmed claude-opus-5-5 (releasedOn 2026-09-22) and claude-fable-5-1 (releasedOn 2026-09-01) directly against the Anthropic docs' embedded page JSON; sharpened the OpenAI entry to state the docs page itself carries no release-date field for gpt-6-sol/gpt-6-luna (only knowledge-cutoff dates), and that 2026-09-22 is CNBC/secondary-source only. Source: verify/gen/anth_models_opus-5-5_overview.html, anth_about-claude_models_overview.html (\"releasedOn\":\"2026-09-22\", \"releasedOn\":\"2026-09-01\"); verify/gen/oai_developers.openai.com_api_docs_models.html.

Mac target: M5 Pro, 64GB unified memory, about 53 GiB Metal-available, budget about 40GB at Q4 including KV for 32K context.
llama.cpp floor: b11057. Diffed against b11214 (2026-09-27T12:45:21Z, newest tag at sweep time); only 2 new HF-arch registrations appeared between them (BailingMoeV3VL, Gemma4DSparkModel), neither relevant here.
Method: HF API createdAt-desc sweep of 18 orgs (plus corrected slugs stepfun-ai, prism-ml, XingChen-AGI), then a second lastModified-desc sweep of the same 18 to catch updates to pre-window repos (44 hits reviewed, listed below).
Evidence labels: [doc]=publisher documentation/card, [obs]=observed this session (HF API/config/tree/Ollama fetch), [self]=publisher-reported benchmark, [vendor-harness]=publisher-run but against third-party baselines, [unk]=not verified this session.
Release date vs createdAt: HF repo createdAt often precedes the publisher-stated public Release Date on the card. Where the card states one, both are given.

Sorted by Mac fit, ascending estimated total footprint within each bucket

Model Org (in 18-list?) Repo created [obs] Card Release Date [doc] Revision (sha, 12-hex) Total / Active params [obs] License [obs+doc] Mac artifact Ollama tag [obs] llama.cpp arch @ b11057 [obs] Q4 weight KV f16 @32k Total est Fit
ibm-granite/granite-4.2-3b ibm-granite (yes) 2026-08-07 unk e459acceac81 3.66B dense apache-2.0 Official GGUF + official MLX granite4.2:3b, 2.2GB GraniteForCausalLM: yes 2.24GB [obs] 2.68GB 6.4GB Fits comfortably
tencent/ContextPilot-E4B tencent (yes) 2026-08-27 unk a7ba41c18ddc 7.94B (Gemma4 MatFormer effective-4B) Tencent custom LICENSE file (HF tag: other) mradermacher GGUF (i1 imatrix) none found Gemma4ForConditionalGeneration: yes 5.30GB [obs] not resolved, nested config approx 8-9GB Fits comfortably
tencent/ContextPilot-8B tencent (yes) 2026-08-27 unk 595abeaed72e 8.19B dense (Qwen3-8B base) Tencent custom LICENSE file (HF tag: other) mradermacher GGUF (i1 imatrix) none found Qwen3ForCausalLM: yes 5.03GB [obs] 4.83GB 11.4GB Fits comfortably
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B XiaomiMiMo (yes) 2026-09-21 unk 2367e865d009 9.41B (Qwen3.5-9B distill) mit bartowski+ggml-org GGUF, mlx-community 4bit none found Qwen3_5ForConditionalGeneration: yes approx 5.6GB (Q4_K_M est; Q8_0 used in baseline test) approx 4.8GB approx 11-12GB Fits comfortably; ALREADY TESTED, see note
ibm-granite/granite-4.2-8b ibm-granite (yes) 2026-08-07 unk f8de16cdcdbc 8.79B dense apache-2.0 Official GGUF + official MLX granite4.2:8b, 5.3GB GraniteForCausalLM: yes 5.35GB [obs] 5.37GB 12.2GB Fits comfortably
tencent/ContextPilot-14B tencent (yes) 2026-08-27 unk 8eafd8356c5c 14.77B dense (Qwen3-14B base) Tencent custom LICENSE file (HF tag: other) mradermacher + bartowski GGUF none found Qwen3ForCausalLM: yes 9.00GB [obs] 5.37GB 15.9GB Fits comfortably
meta-models/Muse-Glimmer-30B (NOT in 18-org list) meta-models 2026-08-09 unk a4e59da52a7b 29.78B (nested config) apache-2.0 unsloth GGUF, mlx-community 8bit none found MuseGlimmerForConditionalGeneration: yes 15.88GB [obs, UD-Q4_K_XL] not resolved approx 18-20GB Fits comfortably; NO benchmark gathered
Qwen/Qwen3.8-27B (REFERENCE: existing default, in-window release) Qwen (yes) 2026-08-05 unk 1d4bf0f2ff60 27.78B, HYBRID attention (16 of 64 layers full-attention, interval 4; rest linear/SSM-style with O(1) state) apache-2.0 Official; unsloth GGUF (proven quant) qwen3.8:27b, 18GB Qwen3_5ForConditionalGeneration: yes (proven) 16.46GB [obs, unsloth UD-Q4_K_M] approx 2.15GB (only 16 full-attn layers, 4 kv heads, head_dim 256; linear layers add a small constant state, not context-scaling) approx 20.1GB Fits comfortably; smallest 32K footprint among rows with a worked-out KV in the approx 27-30B class here (Muse-Glimmer-30B's KV is not worked out, so it is not a counted comparison)
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 nvidia (yes) 2026-08-01 2026-08-11 [doc, card Release Date field] a9904d24bcc1 31.58B total / 3B active, hybrid Mamba+MoE (NemotronH) OpenMDW-1.1 unsloth/bartowski/ggml-org/lmstudio-community GGUF (no official) nemotron-3.5-lightning:30b-a3b, 25GB NemotronHForCausalLM: yes 25.27GB [obs, unsloth UD-Q4_K_M] 0.20GB [obs, config layers_block_type: 6 attention / 23 mamba / 23 moe of 52 layers; 2 kv heads x head_dim 128, only the 6 attention layers scale with context] approx 27.0GB (25.27 + 0.20 + 1.49) Fits comfortably; below granite-4.2-30b's 27.8GB at 32K, but still 6.9GB above Qwen/Qwen3.8-27B's 20.1GB two rows up - not the smallest of the class
ibm-granite/granite-4.2-30b ibm-granite (yes) 2026-08-07 2026-08-25 [doc, card Release Date field] 9e668ce1c538 29.28B dense apache-2.0 Official GGUF + official MLX, also lmstudio-community/bartowski/mradermacher granite4.2:30b, 18GB GraniteForCausalLM: yes 17.72GB [obs] 8.59GB (plain GQA, all 64 layers full-attention; no hybrid discount) 27.8GB Fits comfortably; costs approx 7.7GB MORE than Qwen3.8-27B at 32K despite similar weight size

runtime_unsupported (mainline llama.cpp b11057 or b11214)

Model Org Created [obs] Revision Params License Why unsupported Note
XingChen-AGI/Xing4.0-29B-A4B (NOT in 18-org list) XingChen-AGI 2026-09-16 baae3c3e813c 31.2B total/4B active MoE, full MHA (32 kv=32 attn heads, no GQA) apache-2.0 Xing4_0ForCausalLM not registered at b11057; reconfirmed absent at b11146 and b11214 (2026-09-27T12:45:21Z, the newest tag) via a fresh architecture-registry scan Even if supported, full-MHA 32K KV alone is approx 18.79GB, worst of this table
prism-ml/Ternary-Bonsai-2-27B-gguf (base Qwen/Qwen3.8-27B) prism-ml (closest match to task's PrismML) 2026-09-16 b072e1d3b35a approx 27B ternary, 1.72-2.13 bits/weight apache-2.0 Custom ternary hybrid-attention kernel; own card confirms stock llama.cpp rejects the shipped PQ2_0/PTQ1_0 files outright as unknown types (does not run at all) and silently loads the unrelated generic Q2_0 type with no Hadamard activation runtime, producing garbage output — "produces garbage" applies only to that Q2_0 fallback, not to PQ2_0/PTQ1_0; needs the PrismML-Eng/llama.cpp FORK approx 5.9-7.2GB per card [self]; card measures approx 28 tok/s TG128 on an Apple M5 Pro (this Mac's target chip: 28.1 tok/s PQ2_0 current build, 28.7 tok/s on an earlier pre-rotation 7.2GB build pending re-measurement — readme_bonsai2_27b_gguf.md lines 191/206); a separate figure of approx 47 tok/s is claimed for the wider Apple M5 Max, publisher-claimed and not independently verified

Too large for this Mac regardless of quant (total safetensors element counts [obs]; MoE = total, not active)

New since 2026-08-01: deepseek-ai/DeepSeek-V4-Pro-0813 1650.5B mit (approx 990GB@Q4); XiaomiMiMo/MiMo-V2.6-Pro-RL 1024.2B mit (approx 615GB); deepseek-ai/DeepSeek-V4.1-Flash 763.2B mit (approx 458GB, baseline watch); zai-org/GLM-5.3 753.3B glm-5.3-custom (approx 452GB, baseline watch); tencent/Hy4-preview 780.0B apache-2.0 (approx 468GB, new Hunyuan successor); nvidia/Nemotron-3-Labs-Ultra-Math-SFT+RL and nvidia/NVIDIA-Nemotron-Labs-Teacher x5 (STEM/General-Reasoning/Instruction-Following/Competition-Coding/Chat), all 560.5B, license other (approx 336GB each); XiaomiMiMo/MiMo-V2.6-Flash-RL 310.8B mit (approx 186GB); zai-org/GLM-5.3-Flash 321.3B mit (approx 193GB, baseline watch); deepseek-ai/DeepSeek-V4-Flash-Vision-Exp 304.6B mit (approx 183GB).
Unchanged from baseline watch list, still too large: Qwen/Qwen3.8-2.4T-A95B (2.4T/95B active); moonshotai/Kimi-K3 (2.8T/104B active; lastModified moved to 2026-09-02, no new repo -- card-level change only); Qwen/Qwen3.8-Flash-Next (approx 180B stored/6B active).

Official alt-quants of the existing default (not new models)

  • Qwen/Qwen3.8-27B-FP8 (created 2026-08-13, 27.78B, apache-2.0): official FP8 of the same Qwen3.8-27B; approx 27.8GB at FP8, larger than the proven UD-Q4_K_M. nvidia/Qwen3.8-27B-NVFP4 (created 2026-09-04) is a CUDA-kernel-specific mirror, not a Metal/llama.cpp path.

lastModified sweep (catches update-not-create; 44 in-window lastModified hits reviewed across the 18 orgs, createdAt before 2026-08-01)

No previously-unknown Mac-fit foundation model surfaced this way; all 44 hits were either (a) already-covered too-large siblings getting card/file touches (GLM-5.2, Nemotron-3-Ultra-550B-A55B, Nemotron-3-Super-120B-A12B, Kimi-K2-Thinking-NVFP4, Qwen3.6-35B-A3B-NVFP4, GLM-5.2-NVFP4, DeepSeek-V4-* NVFP4 mirrors), (b) non-generation models (OCR, guardian/safety, parse, calibration, audio), or (c) small older items just touched (Phi-4-reasoning-vision-15B, Ternary-Bonsai-27B-gguf original). Two items worth naming: deepseek-ai/DeepSeek-V4-Flash-0731 was created 2026-07-31 (one day before the window) and only touched 2026-08-01, a boundary case not otherwise counted here. nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning (created April 2026, 30B/3B active, touched 2026-08-24/25) is an older same-size-class sibling of the Lightning-30B-A3B row above.
Cross-reference: the Nemotron-3.5-Lightning card benchmark table names Qwen-3.6-35B-A3B and Nemotron 3 Super as stronger peers and Nemotron 3 Nano / GPT-OSS-20B as weaker; nvidia/NVIDIA-Nemotron-3-Super-120B-A12B (120B/12B active) and nvidia/Qwen3.6-35B-A3B-NVFP4 (mirror) confirm those identities but both are pre-window and/or too large for this Mac at Q4.

Orgs checked with no new generation-model release in the window

Org Newest relevant repo observed Created [obs]
meta-llama Llama-Guard-4-12B 2025-04-23, nothing since
openai (HF weights) gpt-oss-safeguard-20b/120b 2025-09-18
mistralai Magistral-Small-2507-GGUF 2025-07-23
google (generation family) gemma-4 QAT variants 2026-06-05 (timesfm-3.0 2026-08-24 is forecasting, non-commercial, baseline)
baidu ERNIE-4.5-VL-28B-A3B-Thinking 2025-11-07
MiniMaxAI MiniMax-M3 / M3-MXFP8 2026-06-02
microsoft (chat/text LLM) Fara1.5-27B/4B agent VLM 2026-07-17 (Phi newest Phi-Ground-Any 2026-05-07)
allenai OLMo-3 family 2026-02-19 to 02-28 (Sept HF OLMo-3 activity is third-party fine-tunes)
stepfun-ai Step-3.7-Flash family 2026-05-23 to 05-28
PrismML / prism-ml ternary GGUF/MLX/AWQ only, no plain safetensors base newest gguf-dev 2026-09-17

Hosted models (release dates only, per task; primary source fetched 2026-09-27)

  • OpenAI, https://developers.openai.com/api/docs/models (fetched 2026-09-27): gpt-6-astra still the flagship, same model ID, knowledge cutoff Apr 30 2026 shown; the page has no version/snapshot field so unchanged-vs-refreshed cannot be fully confirmed from it alone. Two new siblings are live on the same index: gpt-6-sol (knowledge cutoff Apr 20 2026) and gpt-6-luna (May 18 2026), both absent from the baseline record. The OpenAI docs page itself carries no release-date field for either — only the knowledge-cutoff dates above are on-page; the 2026-09-22 expansion date is secondary-source only (WebSearch, CNBC), not corroborated by developers.openai.com.
  • Anthropic, https://platform.claude.com/docs/en/models/overview and .../opus-5-5/overview (fetched 2026-09-27; releasedOn fields read directly from the pages' embedded JSON, confirmed against the raw page source): claude-opus-5-5, releasedOn 2026-09-22, lifecycle active, latest true -- newer than the baseline-pinned claude-opus-5[1m] (2026-07-24). claude-fable-5-1, releasedOn 2026-09-01, matches baseline exactly, remains user-excluded (not re-argued here). claude-sonnet-5, releasedOn 2026-06-30. The same release wave also names claude-mythos-5-1 (paired with Fable 5.1 in the announcement link); not independently investigated.

Unknowns

  • No independent (non-publisher) benchmark located this session for granite-4.2-30b, Nemotron-3.5-Lightning-30B-A3B, ContextPilot, or Muse-Glimmer-30B; all quality evidence above is publisher self-reported or publisher-run-harness (Nemotron explicitly: NeMo Gym/Evaluator, publisher-run against third-party baselines).
  • ContextPilot-E4B and Muse-Glimmer-30B use nested multimodal configs; exact per-layer KV math not resolved this session.
  • Whether Xing4.0-29B-A4B or Ternary-Bonsai-2-27B work on any mainline llama.cpp branch newer than b11214, or only on forks/custom builds, was not re-verified beyond the architecture-registry check.
  • Exact Q4_K_M file size for the current MiMo-V2.6-Distill-Qwen-9B revision was not refetched (baseline receipt used Q8_0); figure above is an estimate.
  • meta-models (Muse-Glimmer) and XingChen-AGI are outside the mandated 18-org list; included because they surfaced via the nvidia/prism-ml mirror trail.
  • Publisher Release Date fields were found only on the granite-4.2-30b and Nemotron-3.5-Lightning cards; other rows show HF createdAt only, which can precede the public announcement.

Revisions (40-hex), exact repo id : sha

  • deepseek-ai/DeepSeek-V4-Flash-Vision-Exp: 6821d6ad3681a4b137b066b76094fa82ebd0a380
  • deepseek-ai/DeepSeek-V4-Pro-0813: 72e1d3230f6c080a530b0a1d46f8eb4602340597
  • deepseek-ai/DeepSeek-V4.1-Flash: dba1be0a40aa45a94ad051997016db3960a90277
  • ibm-granite/granite-4.2-30b-GGUF: 27b350a791e81d9a4d1ddca1c49282e9ec533768
  • ibm-granite/granite-4.2-30b-q4-mlx: d3f143b440062060eb05f7b8677ced5b4d9a8946
  • ibm-granite/granite-4.2-30b: 9e668ce1c538387ef24d3644e9b0606647762636
  • ibm-granite/granite-4.2-3b-q4-mlx: 0c6f39b1827afd5eb2c1c3b13751929857434953
  • ibm-granite/granite-4.2-3b: e459acceac81e5fe67c07d9cfc72329a332e7eb1
  • ibm-granite/granite-4.2-8b-q4-mlx: 9047f073e46c527f70dade07bb55b1700d728e36
  • ibm-granite/granite-4.2-8b: f8de16cdcdbc6c779ca517604e050d82cc119e44
  • nvidia/Muse-Glimmer-30B-NVFP4: 47818374517751c48c55cde2621594926b1888b6
  • nvidia/Nemotron-3-Labs-Ultra-Math-SFT: c6337c397cc7ea2cf7b594e0b880bc874558e20d
  • nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16: 434456c9a6753f29d24e23c95d622aaf17111b3b
  • nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16: a9904d24bcc1d289a1950fa9d2b978c47cf903b9
  • nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4: bee7596271d1495f6992ae224aefde4410e816b8
  • nvidia/NVIDIA-Nemotron-Labs-Teacher-Chat: 7fbf767583dfd7f18f11734626ce1fb311d6acb1
  • prism-ml/Ternary-Bonsai-2-27B-gguf: b072e1d3b35a0a630cece372c2127528e0994386
  • Qwen/Qwen3.8-27B-FP8: 017b9c7af6b5689d5dd426a76e0bc077eb5ca20a
  • tencent/ContextPilot-14B: 8eafd8356c5c5e713e92e623c8f1d346de6d0f95
  • tencent/ContextPilot-8B: 595abeaed72e4a4d9da94d04917f0618989a5553
  • tencent/ContextPilot-E4B: a7ba41c18ddcd334f23f63ef7a608ca46bfe3aa4
  • tencent/Hy4-preview-FP8: 4215ec29de873a998e849cee902654490c7ff4d1
  • tencent/Hy4-preview: 705d81ee51566a186d645b74c974d642ef2828fe
  • tencent/UI-Mate-27B: 3ade2378fc84032d5017c1a9c93c4eaa77d65e57
  • tencent/UI-Mate-9B: 05dd5f2975195a5bb03d4363e8767f12158c8421
  • XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B: 2367e865d009c13ac81713a2878291d33ab28177
  • XiaomiMiMo/MiMo-V2.6-Flash-RL: 5711b268169967567844e1e560e8a3966da959b1
  • XiaomiMiMo/MiMo-V2.6-Pro-RL: 73875d00b30a89ef8cc353a0b60b0e9f9561952d
  • XingChen-AGI/Xing4.0-29B-A4B: baae3c3e813cad5f888f1f485cfff659c89076c5
  • zai-org/GLM-5.3-Flash: eb9eb208eb0d988989d07a6a12d0fdeb5f52574a
  • zai-org/GLM-5.3: aca966e4e02791568aa6a4ced368624b3d897f42
  • Qwen/Qwen3.8-27B: 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 (base card)
  • unsloth/Qwen3.8-27B-GGUF: 4ca720788d1e01f1bff70c033e0d0028fd02e502 (proven quant, per baseline)
  • meta-models/Muse-Glimmer-30B: a4e59da52a7bc87ae7251dd5545c0dd437c44b68
  • moonshotai/Kimi-K3: f831ab66814297da540d832a5235f8e904f29d06 (per baseline, unchanged)
  • Qwen/Qwen3.8-Flash-Next: de4b8e4d43b917e7706784d8bb445c9af86a3540 (per baseline, unchanged)
  • Qwen/Qwen3.8-2.4T-A95B: 207bd685a7e3696cfaff12ded7c6a7ea0f88c996 (per baseline, unchanged)

Protocol caveats (added after review)

  • granite-4.2-30b README line 347 states thinking is enabled by default; the card benchmarks quoted above (AIME25 89.17, GPQA 66.41, MMLU-Pro 77.60) are THINKING-MODE scores. The baseline proven extraction protocol for Qwen3.8-27B runs thinking DISABLED at about 17 output tokens per filing. These numbers do not transfer to that protocol; granite-4.2-30b has no thinking-off, task-matched evidence.
  • Ternary-Bonsai-2-27B self-reported 84.78 average is likewise across 14 THINKING-MODE benchmarks per its own card; same non-transfer caveat, on top of it being runtime_unsupported on mainline.
  • tencent/ContextPilot-E4B license was read from the sibling ContextPilot-8B repo LICENSE file (Tencent custom terms); E4B itself is Gemma4-based and was not individually read this session. Mark E4B license as [unk: LICENSE file present, not opened] rather than assuming it matches the 8B terms.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

S3 input from mac-coordinator-64gb-20260925, 2026-09-27: memsys sweep table (suggested path: inputs/sweep-20260927-mac/memsys.md). Sonnet sweep, independently verified by an Opus pass against primary sources; its "Verification" section lists every correction applied. The published and self-reported numbers here are discovery evidence, not S3 evidence.

Table

Memory-systems sweep — 2026-09-27

Verification (2026-09-27, independent Opus pass)

  • ai-memory: 19b6429 vs v2.4.1 ahead/behind counts were swapped (correct: ahead_by 46, behind_by 62, merge base 8ee81bede45b); added that v2.4.0 is an ancestor of 19b6429 and v2.4.1 is not, and reframed the rollback-pin move as a trade-off rather than a recommendation. No local evidence file for this repo's git compare exists in this session's verify directory; command: gh api repos/akitaonrails/ai-memory/compare/v2.4.1...19b6429.
  • GBrain: removed "now documented", "weakened" and "now feasible" — the local-provider docs already existed at the baseline pin v0.56.2.0, and embedding/reranking still default to hosted Voyage; #5284/#5449 are still open and #5449's fix series has not shipped in v0.59.0.0. No local evidence file for this repo in this session's verify directory; source as given: github.com/garrytan/gbrain issues #5284, #5449.
  • claude-mem v13.28.0: corrected the installer-default quote — the general interactive (npx) installer pre-selects hosted "CMEM Pro" (a hosted observer, not a sync tier), the non-interactive default is the Anthropic (local) plan, and telemetry is on by default (opt-out). No local evidence file for this repo in this session's verify directory; source as given: thedotmack/claude-mem's installer script/CLI text.
  • mcp-memory-service v11.14.0: confirmed Codex CLI is listed among supported MCP clients (client support, not native hooks). Source: verify/memsys/pypi-mcp-memory-service.json (package description, "CLI & Terminal AI (MCP)" section).
  • Hindsight: added that the newer integration tags (integrations/coding-agents/v0.7.0 @ 0c0869b7, 2026-09-24; paperclip, obsidian, hermes) are not visible through releases/latest, and recommended pinning integrations/coding-agents/v0.7.0 for the capture path; recorded the Mac trial blocker (pg0-embedded 0.15.2 / macOS Postgres 18.1.0 / openssl@3 dylibs / DYLD_* stripped under initdb's popen//bin/sh call). verify/memsys/pypi-hindsight-all.json confirms only the package version (0.10.1); the integration-tag and Mac-blocker specifics have no local evidence file in this session and are taken as given.
  • Cognee: refined the disqualifier correction already in the draft with the specific commit (topoteretes/cognee-integrations @ 4042bcf8, unpinned) and the running-server requirement, and made explicit that it stays deferred for other reasons (Docker/graph-store weight, no published benchmark). No local evidence file for this repo in this session's verify directory; source as given: topoteretes/cognee-integrations.

Baseline: catalogs/foundation/memory-stack-20260925.json (repository_decisions) +
evidence/artifacts/memory-stack-20260925/convergence.json, both in
the repository checkout. All tags/commits below resolved 2026-09-27 via
gh api repos/{owner}/{repo}/releases/latest (falls back to /tags if no Release object
exists — none needed it) then git/ref/tags/{tag} → git/tags/{sha} for annotated tags,
i.e. GitHub Releases only, never a main-branch HEAD. Full commands run are in
gh_fetch.sh / gh_raw.tsv in this directory. Verified byte-identical against the
baseline's own pins for agentmemory, MemPalace, Hindsight, Honcho, OpenViking, Basic
Memory, Letta — the resolution method matches theirs.

Baseline systems

System Latest release @commit, date License Local-on-Mac (no hosted key) Claude/Codex integration Storage/services Best published benchmark 2026-09-25 decision Suggested change
ai-memory (akitaonrails/ai-memory) — production Running: 19b6429 local 2.5-pre build (reports 2.4.0), not an official artifact. Official: v2.4.1 @ 433a19f3d54d, 2026-09-25T21:46 MIT Yes — Ollama 0.34.4 + qwen3-embedding:4b + qwen3.5-9b-64k, no hosted key Production memory server shared by Claude Code and Codex (operational record; hook/plugin detail not re-verified this wave) Ollama local serving, local store NCE (Mac, descriptive): C3 aimem-qwen3 0.570 recall_all@5 (470/470), 20/20 on a drift-check reproduction. evidence/artifacts/memory-stack-20260925/convergence.json#/rows/0 retain (control) Verified: 19b6429 vs v2.4.1 = diverged (ahead_by 46, behind_by 62; merge base 8ee81bede45b) — corrects an earlier swap of these two counts. v2.4.0 is a strict ancestor of the running 19b6429; v2.4.1 is not. v2.4.1 is a 2.4.x maintenance release (Windows checksum LF fix, +126 commits over v2.4.0), NOT the running 2.5 code and NOT "the first official 2.5 release" the next-test needs. No official 2.5 exists yet. Trade-off, not a recommendation: moving the rollback pin to v2.4.1 means falling back to 62 commits of code production never ran, while also losing the 46 commits production did run that v2.4.1 lacks; staying on v2.4.0 loses only the latter, since it is a pure ancestor and introduces nothing untested. Decision (retain, control) unchanged; the rollback-pin choice is left open as a trade-off, not resolved here.
agentmemory (rohitg00/agentmemory) — trial v0.9.29 @ 2d38dafede67, 2026-08-16T13:56 — unchanged Apache-2.0 Yes — keyless mode (am-keyless) or MiniLM in-process README confirms: native plugin + 12 hooks + 54 MCP tools (Claude Code); native plugin + 6 hooks + MCP (Codex CLI); also Cursor/Gemini CLI/OpenCode/etc. Note: baseline's Codex caveat (openai/codex#16430) is about Codex Desktop, a different surface from the Codex CLI hooks this README documents — re-verify which surface D2h targets Local memory server on :3111, REST+MCP, no external DB required NCE (Mac): am-minilm 0.821 recall_all@5 (470/470), +25.1pp over ai-memory C3 (Holm p=0.0001). Own bench doc: github.com/rohitg00/agentmemory/blob/main/benchmark/LONGMEMEVAL.md. convergence.json#/rows/1 trial None — main moved (pushed 09-26) but no new tag. Clarify Codex CLI vs Desktop hook-dispatch scope before D2h.
MemPalace (MemPalace/mempalace) — trial v3.10.0 @ 22fd87f09c19, 2026-09-16T00:24 — unchanged MIT Yes, local palace mode Hooks for both Claude Code and Codex + MCP (per baseline); not independently re-verified this wave Local palace store; optional ghcr container DV/vendor-scored: recall_all@5 0.857 raw / 0.887 hybrid v4 (BENCHMARKS.md 2026-04); third-party arXiv 2604.21284 puts palace modes at 89.4%/84.2% recall_any. Not independently reproduced by us. trial None — main moved (pushed 09-25), no new tag.
Hindsight (vectorize-io/hindsight) — trial v0.10.1 @ f8950b0c07d9, 2026-09-21T15:24 — unchanged MIT Yes, but LLM-extraction-on-every-retain is heavy (~86h/470q at gpt-oss-20b Mac rate) Shared capture for Claude and Codex via official coding-agents package (per baseline); newer integration tags are not visible through releases/latest: integrations/coding-agents/v0.7.0 @ 0c0869b7, 2026-09-24, plus separate paperclip/obsidian/hermes integration tags Postgres + worker + local LLM DV: LongMemEval-S QA 91.4% (Gemini-3) / 83.6% (gpt-oss-20b), arXiv 2512.12818 — reproduced only by co-authors, not independent. BEAM-10M 64.1%. trial Pin integrations/coding-agents/v0.7.0 for the capture path (not surfaced by releases/latest). Mac trial blocker recorded: pg0-embedded 0.15.2's macOS Postgres 18.1.0 links /opt/homebrew/opt/openssl@3 dylibs; initdb runs postgres -V through popen//bin/sh, so macOS strips DYLD_* and the embedded DB cannot start on a Mac without Homebrew.
Attemory (AttemorySystem/attemory) — defer v0.1.3 @ 603c03afa9a0, 2026-08-13T06:54 — unchanged MIT (repo); prebuilt core SDK unverified Unverified — core is a closed prebuilt SDK MCP support "only planned" (baseline) Unverified (closed core) DV: 96.38% recall_all@5 on cleaned LME-S (local Qwen3.5-9B, undated) defer None found; dormant since release (pushed_at 2026-08-13, no activity since).
GBrain (garrytan/gbrain) — defer v0.59.0.0 @ e78f1c38b947, 2026-09-26T23:00 — changed (was v0.56.2.0 @ db56c778; 3 minor bumps in ~48h) MIT Local-provider docs (Ollama / llama.cpp llama-server / LM Studio embeddings + a fully-local llama-server Qwen3-Reranker recipe) already existed at the baseline pin v0.56.2.0 — not new this wave. New-install default, and embedding/reranking generally, still default to hosted Voyage voyage-4. Not confirmed as a native Claude Code/Codex plugin this wave (generic CLI/API + docs/integrations/) — unknown PGLite (embedded) DV: v0.59.0.0 changelog reports an "evidence-checking" QA reader moving 308→324/361 judged-correct (n=361, not recall_all@5/470 — different metric, not comparable). Prior: 95.53% "sessions within 5 chunks" (chunks ≠ sessions), 86.6% QA, hosted models. defer "Relies on hosted models" stands: embedding and reranking still default to hosted Voyage, and the local-provider docs cited above already existed at the baseline pin, not a new capability. reindex/embed hangs remain open: #5284 (PGLite COMMIT wedge, updated 09-23) and #5449 (checkpoint self-deadlock, opened 09-24) — #5449 has a fix series but it has not shipped in v0.59.0.0. Defer stands.
Mem0 OSS (mem0ai/mem0) — defer Python OSS v2.2.1 @ 94c3fe9f238f, 2026-09-25T17:35 — changed (was v2.2.0 @ 47a69e1e72dc; +15 commits/47 files). Repo-wide "latest release" is ts-v3.3.1 (TS SDK, same commit, published 75s later) — different package line, not the baseline's Python pin. Apache-2.0 Needs glue code; v2.2.1 fix #7350 stops ConfigManager injecting OpenAI's baseURL/model into other providers — small win for local/non-OpenAI use Claude Code plugin needs a Platform key (not passwordless) Pluggable vector store, Python/TS SDKs DV: managed platform 94.4% LongMemEval / 92.5% LoCoMo / 64.1% BEAM-1M / 48.6% BEAM-10M (mem0.ai/blog, updated 2026-09-22, self-reported); OSS 91.0%/88.6%; independent-ish re-run 73.8% (Maximem, 2026-05-27) defer None material; patch available (v2.2.1), includes a local-provider-relevant fix.
Graphiti/Zep (getzep/graphiti) — defer v0.30.2 @ eaa4128681bc (now resolved; baseline had it unresolved), 2026-09-08T20:38 — unchanged Apache-2.0 No — MCP server needs Docker + Neo4j Generic MCP server only; no Claude-Code- or Codex-specific plugin found in README Neo4j + Docker Zep 71.2% LongMemEval (GPT-4o judge) per mem0.ai/blog (2026-09-22), "consistent across Zep's own reporting and an independent academic comparison" per that source, not independently confirmed by us; LoCoMo 94.7% claimed by Zep vs. 75.1% in ByteRover's third-party re-test (disputed) defer None; Neo4j weight confirmed.
Supermemory (supermemoryai/supermemory) — defer server-v0.0.8 @ 5d2b5855fe49 (now resolved), 2026-08-17T18:39 — unchanged MIT (repo); server ships as binaries No for the documented/default path — primary MCP endpoint is hosted (mcp.supermemory.ai); self-hosting the binary server is possible but not the advertised path Has an explicit Claude Code plugin (github.com/supermemoryai/claude-supermemory) + hosted MCP for many clients Hosted service by default; binary self-host alternative No LongMemEval/LoCoMo/BEAM score found this wave — unknown defer None; flag hosted-by-default MCP as a local-on-Mac concern.
Cognee (topoteretes/cognee) — defer v1.6.1 @ eb90d0374075 (now resolved), 2026-09-24T17:50 — unchanged Apache-2.0 Needs Docker for its MCP/UI service; supports local LLM/embedding providers via .env (not confirmed no-hosted-key by default) Corrects the baseline's group disqualifier: README shows both a Claude Code plugin and a Codex plugin with hooks (~/.codex/config.toml → hooks = true, codex plugin add cognee@cognee), both in topoteretes/cognee-integrations @ 4042bcf8 (unpinned); both need a running Cognee server Docker; graph store + vector store No LongMemEval/LoCoMo/BEAM score found this wave — unknown defer Disqualifier correction: the "no Claude/Codex plugin" disqualifier no longer applies — cognee has native Claude Code + Codex plugins with hooks (topoteretes/cognee-integrations @ 4042bcf8, unpinned, needs a running Cognee server). It stays deferred for other reasons: Docker/graph-store weight, no published retrieval benchmark.
claude-mem (thedotmack/claude-mem) — defer v13.28.0 @ 57c383037344, 2026-09-26T18:37 — changed (was v13.25.3; 5 releases in ~24h: 13.26.0 cloud-sync migration off Cloudflare → new sync.cmem.ai hub, 13.26.1 sync fixes, 13.27.0 agent-cost-report skill pricing Claude Code + Codex transcripts via OpenRouter, 13.27.1 fixes, 13.28.0 non-interactive installer defaults to Anthropic provider) Apache-2.0 Core plugin is local (SQLite + hook scripts). The general interactive installer (npx) pre-selects hosted "CMEM Pro" — a hosted observer (inference via cmem.ai), not a sync tier as an earlier draft called it; the non-interactive installer default is the Anthropic plan (local memory). Telemetry is on by default (opt-out). Native Claude Code plugin: 5 lifecycle hooks (SessionStart, UserPromptSubmit, PostToolUse, Stop, SessionEnd) + MCP search tools. No evidence of a native Codex lifecycle hook — it only reads Codex transcripts for the cost-report skill. Local SQLite + worker; CMEM Pro is a hosted observer (inference via cmem.ai), pre-selected by the interactive installer, opt-out via --provider host or the non-interactive install path No LongMemEval/LoCoMo/BEAM score found this wave — likely the actual disqualifier vs. the group's generic text defer None to the decision; the CMEM Pro hosted-observer surface should be checked for "local-on-Mac, no hosted key" before any trial, since the general interactive installer pre-selects it.
MemOS (MemTensor/MemOS) — defer v2.0.34 @ 41bf5c7fa89e (now resolved), 2026-09-23T05:24 — unchanged Apache-2.0 Not re-verified this wave Not re-verified this wave Not re-verified this wave DV only, or none (baseline) defer None found; not deeply re-checked beyond version/license this wave.
Honcho (plastic-labs/honcho) — reject v3.2.1 @ 79cb31645f4e, 2026-09-24T19:43 — unchanged, byte-identical to baseline AGPL-3.0 Needs Postgres + Redis + worker; embedder can silently fall back to OpenAI (#915) SDK-based, not hook/plugin native Postgres + Redis + worker DV 90.4% QA reject None.
OpenViking (volcengine/OpenViking) — reject v0.4.21 @ 3fca2577520f, 2026-09-20T08:45 — unchanged, byte-identical to baseline AGPL-3.0 Not re-verified this wave Not re-verified this wave Not re-verified this wave DV LoCoMo 80.32% reject None.
Basic Memory (basicmachines-co/basic-memory) — reject v0.23.2 @ c0bd87c6d5a4, 2026-08-25T20:44 — unchanged, byte-identical to baseline AGPL-3.0 Not re-verified this wave Not re-verified this wave Not re-verified this wave DV LoCoMo R@5 76.4% reject None.
Letta server (letta-ai/letta) — reject (retired) 0.16.8 @ 1131535716e8, 2026-05-14T17:14 — unchanged (last server release; still no server release in >4 months, confirms "retired") Apache-2.0 n/a n/a n/a none published reject None.

New / newly-relevant entrants (released or pushed since 2026-08-01, not in baseline)

System Latest release @commit, date License Local-on-Mac Claude/Codex integration Best evidence 2026-09-25 decision Suggested
Memorix (AVIDS2/memorix) v1.9.6 @ e44f2a296b5f, 2026-09-23T07:57 Apache-2.0 Unverified — "memory living under the Git project" suggests local files; embedding/LLM requirements not confirmed Explicitly markets MCP compatibility with Claude Code, Codex, Cursor, Windsurf, Gemini CLI, Antigravity, OpenClaw, Hermes, Pi, Copilot, Kiro, OpenCode, Trae None found this wave — no LongMemEval/LoCoMo/BEAM score n/a (new) Candidate for defer: gap = no published retrieval benchmark. Re-check next wave.
mcp-memory-service (doobidoo/mcp-memory-service) v11.14.0 @ bed24ef72141, 2026-09-25T09:44 Apache-2.0 Plausibly local (REST API + knowledge graph + "autonomous consolidation"; typical of this class) but not independently confirmed this wave Markets itself for Claude + LangGraph/CrewAI/AutoGen pipelines; lists Codex CLI among its supported MCP clients (a compatible-clients listing — client support, not a native Codex lifecycle-hook integration) None found this wave — no LongMemEval/LoCoMo/BEAM score n/a (new) Candidate for defer: same benchmark gap. Re-check next wave.

Considered, not given a row

  • ByteRover / campfirein/byterover-cli (formerly Cipher, 4,966 stars): last release v3.16.1 (2026-05-27), last push 2026-06-25 — outside the since-2026-08-01 window despite appearing at 92.2–96.1% LoCoMo / 92.8% LongMemEval-S in third-party leaderboards (self-reported; mem0.ai's own blog flags internal inconsistency — ByteRover's own comparison table scores Zep at 75.1% and Mem0 at 66.9%, both far below those vendors' self-reports, and ByteRover's 96.1% LoCoMo figure is flagged as identical to a separate "ZeroMemory" claim). License shows NOASSERTION via GitHub API. Has a Claude plugin (campfirein/brv-claude-plugin), no Codex plugin found.
  • JordanMcCann/agentmemory: name collision with the baseline trial candidate (rohitg00/agentmemory) — a different, unrelated repo. MIT license, claims "Fix native worker timeout cleanup on macOS #1 on LongMemEval — 96.2% (481/500)," but created and pushed exactly once on 2026-03-26, zero GitHub Releases, 47 stars — dormant, self-reported, unreproduced. Excluded as neither released nor updated since 2026-08-01.
  • Mastra (mastra-ai/mastra): very active TS agent framework (pushed 2026-09-27), not a Claude-Code/Codex memory plugin itself. Its "Observational Memory" research claims 94.87% LongMemEval (gpt-5-mini judge; benchmark code at github.com/mastra-ai/mastra/tree/main/explorations/longmemeval) per mastra.ai/research. License shows NOASSERTION via GitHub API — unresolved, not manually checked this wave.
  • Agent Zero Memory: arXiv 2608.29606 claims 95.60% LongMemEval / 93.60% LoCoMo (ahead of Mastra 94.87%, Mem0 92.50%, ByteRover 2.0 92.20%) per mem0.ai's blog summary. No discoverable public GitHub repository (searched; only unrelated small hobby projects matched). Paper-only.
  • DolphinBench: a new Mem0-authored memory benchmark ("Mapping the Pareto Frontier of Agent Memory," blog dated 2026-09-22) — not one of the four benchmarks named in this task (LongMemEval-S/M, LoCoMo, BEAM, MemoryAgentBench). Flagged as an emerging benchmark to watch, not scored here.

MemoryAgentBench

No system checked in this sweep publishes a MemoryAgentBench score as of 2026-09-27 (baseline or new entrants). Gap, not resolved.

Scout and others added 2 commits September 27, 2026 10:10
…puts/ (not frozen)

S3 section 2 names the Mac session's 2026-09-27 discovery sweeps as inputs:
embedders, rerankers, memory LLMs, generation models and memory systems,
each Opus-verified with a Verification section listing every correction
and its source. They are committed verbatim from the PR #390 comments,
each with a provenance line (comment URL, time, host), under
blueprints/memory-layer-s3/inputs/sweep-20260927-mac/. A privacy scan found
no personal paths, emails or tokens. The files are registered in
manifests/evidence.json.

They are discovery inputs only: vendor-reported numbers are motivation for
the candidate list, never S3 evidence (section 8). validate.py passed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
…rozen)

GPT-6 accepted r5 (65fa4fe): all six r4 findings resolved, cluster
membership independently matched. Two corrections, neither changing the
protocol:
- the official track excludes 51 of the 56 single-session-assistant
  questions and 5 remain (the frozen eligible manifest, checked against
  the pinned dataset); "the 51" implied all of them;
- section 2 records that the Mac discovery inputs are committed
  (ff12bf1), not pending.
validate.py passed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

S3 text review: GPT-6 accept at r5 (65fa4fee). This was the scoped re-check of r4's six findings; its text follows verbatim. Since then there are two wording corrections (75e30ea), neither changing the protocol:

  • the official track excludes 51 of the 56 single-session-assistant questions, and 5 remain;
  • the Mac inputs are committed (ff12bf1).

Next: build the isolated S3 runner and adapters, then the complete-bundle review (section 9), then the freeze.


VERDICT: accept

All six repairs preserve the accepted replacements’ meaning without weakening them.

  1. High 1 — resolved. Section 1, line 39 specifies native unprefixed input, distinguishes S3 controls from A-protocol reproductions, and requires effective request-byte/configuration verification before scoring. Section 12 and A17:20 agree.

  2. High 2 — resolved. Section 3, line 138 requires canonical (role, content) clustering, excludes IDs and annotations from hashes, and freezes the manifest before scoring. Both stated counts reproduce exactly.

  3. Medium 3 — resolved. Section 3, line 125 separates upstream user-text/relabelled-gold construction from the local full-session/raw-answer_session_ids extension. Both retain unchanged upstream scoring.

  4. Medium 4 — resolved. Section 2, line 79 explicitly selects local MiniLM as a default-provider exception. Shipped hooks or demonstrated equivalence—including SessionEnd transcripts—are required; D2h replay alone is insufficient.

  5. Medium 5 — resolved. Section 7, line 272 propagates excessive fallbacks to every dependent replacement comparison, leaves selection unresolved, and prohibits post-scoring substitution of C3′. Section 6’s preflight substitution occurs before scoring and remains consistent.

  6. Low 6 — resolved. Section 2, line 92 makes discovery inputs a future bundle requirement with dates and hashes. Section 5, line 226 correctly identifies issue [mac-coordinator] model hosting: qualify Mac-native memory/RAG model hosting on the 64 GB Mac #379 and requires its resulting measurement receipt.

No new findings within scope. The repaired definitions, preflight, fallback rule and supersession amendment are consistent.

Verified

Executed the unchanged frozen v4 clusters function. An independent connected-component computation matched every cluster’s membership.

Track Questions Content clusters Singletons + pairs Session-ID clusters
Full 470 452 434 + 18 465
Official 419 401 383 + 18 414

The downloaded pinned dataset matched SHA-256 d6f21ea9d60a0d56f34a05b609c79c88a451d2ae03597821ea3d5a9678c3a442. Counts use the frozen eligibility manifest; the official track retains five assistant-type questions, so excluding that entire question type would be incorrect.

python3 scripts/validate.py passed, exit 0: 69 components, 7,322 hashed files, 4 profiles, 159 receipts. Head is 65fa4fee; checkout remains clean. No edits, commits or posts.

🤖 Generated with Claude Code

https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5

Scout and others added 4 commits September 27, 2026 15:08
…ss-family review (not frozen)

The r6 draft adds the upstream evaluation harnesses (AMB unmodified with frozen
MemoryProvider modules, the unchanged LongMemEval eval_utils scorer, scipy and
statsmodels statistics, MTEB) and the gateway call envelope for GPT-6 arms.
The repair applies all 13 findings of the Claude review (H1-H4, M1-M5, L1-L4):
immutable committed input copies under inputs/r6/, K-concurrency with a 96 h
coordinator-set window (user may overturn before freeze), the latency preflight,
and a new section 15 mapping each finding to its fix. Still DRAFT r6.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Hot-file protocol (docs/lanes.md): main's manifests/evidence.json, the branch's
files re-registered, including the five r6 input copies, and both generated
reports rewritten with their --write commands. validate.py passed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…AMB's retrieval mode; not frozen)

A one-round Claude re-check of the GPT-6 repair found two high defects in the
AMB text the coordinator had directed the repair to adopt, verified here in AMB
03c1d0f1: EvalRunner always constructs a Gemini judge and calls it for AMB's
task_type="open" LongMemEval dataset (runner.py:42, longmemeval.py:67), and
AMB's loader passes gold-revealing "{question_id}_{session_id}" IDs to every
provider (longmemeval.py:328-350). Coordinator decision: one frozen, hashed
LongMemEval dataset adapter passed to unmodified EvalRunner.run in AMB's
retrieval mode (task_type="retrieval": no judge call, runner.py:222-227; opaque
IDs; score_retrieval as the ranked-ID export point), following AMB's own
PrecisionMemBench pattern. Also: cognee v1.6.1 environment versus AMB's 0.5.4
lock, a dispatch wrapper for K concurrency, S3 answerer and judge outside AMB,
validate.py on the pushed tree, attribution and locator precision. Section 16
maps N1-N10; the fixes are not yet independently re-reviewed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Hot-file protocol (docs/lanes.md): main's manifests/evidence.json, the branch's
files re-registered, including the five r6 input copies, and both generated
reports rewritten with their --write commands. validate.py passed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins seathatflowsinourveins changed the title Memory layer S3: head-to-head preregistration draft r4 (governs both hosts; A17 supersession proposed; not frozen) Memory layer S3: head-to-head preregistration draft r6 (both hosts; upstream AMB harness in retrieval mode; not frozen) Sep 27, 2026
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Superseded by #526 (S3 r7): maintenance gate (90-day Scorecard rule; LongMemEval/LoCoMo repos now stale and cited only), current candidate pins, symmetric merit rule (ai-memory v2.4.2 as a contender with no protected status, per the user's 2026-09-30 decision), AMB byte-identical at 03c1d0f1 with thin providers, token ledger and RAG head-to-head. #526 carries this branch's history plus origin/main.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Closing as superseded by #526, as its owner declared (#390 (comment)); #526 retains all 14 changed paths. 2026-10-04 superseded-work cleanup. The branch stays available for history.

🤖 Generated with Claude Code

@seathatflowsinourveins
seathatflowsinourveins deleted the claude/memory-layer-s3-prereg-20260927 branch October 8, 2026 07:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant