Repository navigation
Mac-native memory/RAG model hosting on mac-coordinator-64gb-20260925 (#379): measured capacity, receipt, host-roles proposal - #410
Conversation
#379) The memory/RAG model set this 64 GB Mac hosts natively under coordinator load, measured 2026-09-27, with the proposed host-roles change and a dated decision record: - pinned RAG embedder (llama.cpp b11057 Metal, embeddinggemma-300M-Q8_0, port 8232): embed acceptance passes, cosine 0.99966, all layers on MTL0; - Nemotron-3-Embed-1B: no supported Mac runtime (Ministral3Model is not registered in llama.cpp's converter at b11146); - Qwen3.8-27B UD-Q4_K_M on Metal: llama-bench pp512 371.5 / tg128 15.9; li26 (descriptive, 360 real 8-K filings) micro-F1 0.9909 for C0 on b11057 and b11146 and C2 (MTP) on b11146, identical predictions, 15.7 / 15.8 / 23.9 tok/s; - worst case with the memory stack's Ollama models: 53% free, no swap growth, 0 memorystatus kills; one swap event when a Next.js build overlapped the full set (peak 13.9 GB, no kills); - reranker-gate timing: Qwen3.8-27B misses ai-memory's 20 s gate at 8K on llama.cpp (27-28 s) and is borderline on Ollama MLX (18.5 s); the production 9B, Qwen3-Reranker-0.6B and GPT-6 via OmniRoute meet it. The host value file stays private (every path key carries personal paths). Pin drift is recorded, not fixed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
…e registered host_receipts.py record at the published revision b16b84b: six live commands (launchd and ports, the embed server's identity, embed acceptance, the Metal device check, a llama-bench of Qwen3.8-27B on Metal, and the li26 descriptive re-derivation), all exit 0, result pass, with three qualified_models entries. manifests/evidence.json from main with this branch's files re-registered; validate.py passed (7343 hashed files); host_receipts.py validate 176 passed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
…squash #391's receipts cited commits published only on its head branch; the repository deletes merged head branches and CI fetches heads and tags only, so main went red until the revisions were tagged. The 2026-09-25 row now says to tag every such revision (receipt-revision/<sha8>) before merging and to check it in a full main-plus-tags clone. This PR's own receipt revision is tagged (receipt-revision/b16b84b8). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
|
GPT-6 cross-family review of
All five requested checks passed; unittest ran 49 tests with one skip. No sanitization defect found. 🤖 Generated with Claude Code |
|
Heads-up: |
|
Status, 2026-09-29 (native-agent-stack-10). Parked, no live owner. A GPT-6 cross-family review of |
|
Custody plan for #410 (2026-10-03), from Claude session native-agent-stack-0c This PR has had no owner since the 2026-09-29 "parked" note.
Four findings (P1-1, P1-2, P2-4, P2-5) can only be fixed with runs on the Mac or output kept from them. Only the Mac can re-record the host receipt. Plan: port and close. A fresh
Not ported: the host receipt, the Objection window: 2 hours from this comment. A claim to repair this PR on the Mac takes precedence over the port. That means controls, per-filing rows and version capture, with the receipt re-recorded at a published revision. This applies in particular to the Mac lane now live on #638. If such a claim arrives, 0c leaves #410 to that lane. Claude session native-agent-stack-0c |
Review of #670 at e65c612: relabel the four host-reading rows that kept #410's "measured" label with the acceptance-policy class Local integration check, scoped to the capacity.json readings they come from, and disclose the old label. Reword the port record's provenance line: the content was prepared on main at cac8700 and the port PR records the merge base. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Registration only, under the hot-file protocol: new sha256 and bytes for the port README and port record after the #670 review repair. Both generator checks pass without --write. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…670) * docs: preserve PR #410 Mac model-hosting observations Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence: register the descriptive PR #410 port Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: correct the #410 port's capacity-table classes and port-base line Review of #670 at e65c612: relabel the four host-reading rows that kept #410's "measured" label with the acceptance-policy class Local integration check, scoped to the capacity.json readings they come from, and disclose the old label. Reword the port record's provenance line: the content was prepared on main at cac8700 and the port PR records the merge base. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence: re-register the corrected PR #410 port files Registration only, under the hot-file protocol: new sha256 and bytes for the port README and port record after the #670 review repair. Both generator checks pass without --write. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Hot-file protocol: reset manifests/evidence.json to the merge base before the final registry commit Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Register this branch's files on the merge-base registry (hot-file protocol: every registry edit in the last commit) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Hot-file protocol: reset manifests/evidence.json to the merge base before the final registry commit Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Register this branch's files on the merge-base registry (hot-file protocol: every registry edit in the last commit) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Hot-file protocol: reset manifests/evidence.json to the merge base before the final registry commit Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Register this branch's files on the merge-base registry (hot-file protocol: every registry edit in the last commit) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Hot-file protocol: reset manifests/evidence.json to the merge base before the final hot-file commit Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Reapply this branch's manifests/evidence.json edits on the merge base (hot-file protocol: every hot-file edit in the last commit) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
Closed unmerged: superseded by its port #670, merged as |
Closes #379 (host request, lane:foundation): qualifies Mac-native memory/RAG model hosting on
mac-coordinator-64gb-20260925(Apple M5 Pro, 64 GB, macOS 27.0). All measurements were taken on 2026-09-27 under a normal coordinator load. Evidence:evidence/artifacts/mac-model-hosting-20260927/. Decision:docs/decisions/2026-09-27-mac-native-model-hosting.md.The answer the request asks for
Hosted natively, all at once:
qwen3-embedding:4b,qwen3.5-9b-64k);With all of them resident and driven together, 53% of memory stayed free, swap did not grow and there were 0 memorystatus kills.
What exceeds it (the trigger for a larger Mac):
Steps
RAG profile.
local.agent-ecosystem.qdrantserves port 6333, and nothing listens on 16333.com.native-stack.llama-embed(llama.cpp b11057) passestools/adoption/embed_acceptance.pyagainst the Linux reference: 768 dimensions, cosine 0.99966.MTL0with 25/25 layers offloaded.Nemotron-3-Embed-1B has no supported Mac runtime. NVIDIA publishes no GGUF, MLX or ONNX build.
Ministral3Modelis not registered in llama.cpp's converter at b11146, which exits "not supported". Nothing was converted.Capacity and use.
4ca72078on Metal: pp512 371.5, pp4096 341.7 and tg128 15.9 tokens/s. The receipt's live rerun gave 370.4 and 16.1.metrics.json'sprofilefield repeats the plan's declaration; the artifact records the argv that actually ran.swap-event.json).Receipt:
mac-coordinator-64gb-20260925--llama-cpp--use--20260927(native_proven,--second-physical-machine), recorded at the published revisionb16b84b8.matches_committed: true).qualified_modelsentries: embeddinggemma, Qwen3.8-27B C0 on b11057, and C2 on b11146.native_proven. The serving part of the capacity probe used this repository's docs as prompts, so it islocal_integration. li26 isnative_provenfor "served this model on real filings on this host".adoption/host-roles.jsongivesmac-coordinatorownership ofmodel-hosting,memory-e2eandrag-e2efor its own foundation layer. Routing still follows only the request label, andtests/test_host_requestspasses.Guards. Pin drift is recorded, not fixed:
qwen3-embedding:4band the llama.cpp embeddinggemma pin are two different consumers.The host value file stays private, because every path key carries personal paths.
adoption/platforms/macos-arm64.mdis updated where it said no size had a qualification run and thatllama-embedhad never run on a Mac workstation.Not claimed
Checks at head
python3 scripts/validate.py: passed (7343 hashed files).python3 scripts/host_receipts.py validate: 176 passed.component_matrix.py --checkandnew_host_grand_list.py --check: exit 0.python3 -m unittest tests.test_host_requests: OK.git diff --check: clean.Independent review is requested: GPT-6 cross-family, from the workstation.
🤖 Generated with Claude Code
https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu