Skip to content

Mac-native memory/RAG model hosting on mac-coordinator-64gb-20260925 (#379): measured capacity, receipt, host-roles proposal - #410

Closed
seathatflowsinourveins wants to merge 4 commits into
mainfrom
claude/mac-model-hosting-20260927
Closed

seathatflowsinourveins wants to merge 4 commits into
mainfrom
claude/mac-model-hosting-20260927

Conversation

@seathatflowsinourveins

Copy link
Copy Markdown
Owner

Closes #379 (host request, lane:foundation): qualifies Mac-native memory/RAG model hosting on mac-coordinator-64gb-20260925 (Apple M5 Pro, 64 GB, macOS 27.0). All measurements were taken on 2026-09-27 under a normal coordinator load. Evidence: evidence/artifacts/mac-model-hosting-20260927/. Decision: docs/decisions/2026-09-27-mac-native-model-hosting.md.

The answer the request asks for

Hosted natively, all at once:

  • the pinned RAG embedder (embeddinggemma-300M-Q8_0 on llama.cpp Metal, port 8232);
  • the memory stack's Ollama models (qwen3-embedding:4b, qwen3.5-9b-64k);
  • Qdrant;
  • one 27B-class 4-bit generation model (Qwen3.8-27B UD-Q4_K_M, 19.4 GiB peak at a 32K context).

With all of them resident and driven together, 53% of memory stayed free, swap did not grow and there were 0 memorystatus kills.

What exceeds it (the trigger for a larger Mac):

  • A 27B-class reranker LLM at 8K tokens. Qwen3.8-27B takes 27-28 s on llama.cpp Metal and 18.5 s median on Ollama MLX, against ai-memory's 20 s gate.
  • A heavy build overlapping the full model set. One Next.js production build did this: swap peaked at 13.9 GB, with no kills.
  • Not a memory limit: Nemotron-3-Embed has no supported Mac runtime.

Steps

  1. RAG profile.

    • The guard ran before any launchd step. No second Qdrant or ai-memory was started; the pre-existing local.agent-ecosystem.qdrant serves port 6333, and nothing listens on 16333.
    • com.native-stack.llama-embed (llama.cpp b11057) passes tools/adoption/embed_acceptance.py against the Linux reference: 768 dimensions, cosine 0.99966.
    • llama-server's own log shows MTL0 with 25/25 layers offloaded.
  2. Nemotron-3-Embed-1B has no supported Mac runtime. NVIDIA publishes no GGUF, MLX or ONNX build. Ministral3Model is not registered in llama.cpp's converter at b11146, which exits "not supported". Nothing was converted.

  3. Capacity and use.

    • llama-bench, Qwen3.8-27B UD-Q4_K_M @4ca72078 on Metal: pp512 371.5, pp4096 341.7 and tg128 15.9 tokens/s. The receipt's live rerun gave 370.4 and 16.1.
    • li26, descriptive, on 360 real SEC 8-K filings (not parity): micro-F1 0.9909 and JSON-valid 1.0 for C0 on b11057, C0 on b11146, and C2 (MTP drafting) on b11146. Predictions are identical on all 360 filings. Median decode was 15.7, 15.8 and 23.9 tokens/s.
      • metrics.json's profile field repeats the plan's declaration; the artifact records the argv that actually ran.
    • Worst-case co-residency and reranker-gate timing ran on the newest build, b11214. GPT-6 through this Mac's OmniRoute gateway answers rerank-shaped 8K prompts in 4.7-9.6 s.
    • The swap event is attributed from the minute samples and file birth times (swap-event.json).
  4. Receipt: mac-coordinator-64gb-20260925--llama-cpp--use--20260927 (native_proven, --second-physical-machine), recorded at the published revision b16b84b8.

    • Six live commands, all exit 0: service and port state, the embed server's identity and digest, embed acceptance, the Metal device check, a llama-bench of the 27B, and the li26 re-derivation (matches_committed: true).
    • Three qualified_models entries: embeddinggemma, Qwen3.8-27B C0 on b11057, and C2 on b11146.
    • Evidence classes: llama-bench and embed acceptance are native_proven. The serving part of the capacity probe used this repository's docs as prompts, so it is local_integration. li26 is native_proven for "served this model on real filings on this host".
    • The proposal: adoption/host-roles.json gives mac-coordinator ownership of model-hosting, memory-e2e and rag-e2e for its own foundation layer. Routing still follows only the request label, and tests/test_host_requests passes.
  5. Guards. Pin drift is recorded, not fixed:

    • llama.cpp: b11057 cannot run the production C2 profile, while b11146 and b11214 can;
    • codex: 0.155.1 / 0.157.1 / 0.158.0-alpha.2.1;
    • ai-memory: 2.3.2 / 2.4.1 / the 19b6429 build;
    • embeddings: Ollama qwen3-embedding:4b and the llama.cpp embeddinggemma pin are two different consumers.

    The host value file stays private, because every path key carries personal paths.

adoption/platforms/macos-arm64.md is updated where it said no size had a qualification run and that llama-embed had never run on a Mac workstation.

Not claimed

Checks at head

  • python3 scripts/validate.py: passed (7343 hashed files).
  • python3 scripts/host_receipts.py validate: 176 passed.
  • component_matrix.py --check and new_host_grand_list.py --check: exit 0.
  • python3 -m unittest tests.test_host_requests: OK.
  • git diff --check: clean.

Independent review is requested: GPT-6 cross-family, from the workstation.

🤖 Generated with Claude Code

https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu

#379)

The memory/RAG model set this 64 GB Mac hosts natively under coordinator
load, measured 2026-09-27, with the proposed host-roles change and a dated
decision record:
- pinned RAG embedder (llama.cpp b11057 Metal, embeddinggemma-300M-Q8_0,
  port 8232): embed acceptance passes, cosine 0.99966, all layers on MTL0;
- Nemotron-3-Embed-1B: no supported Mac runtime (Ministral3Model is not
  registered in llama.cpp's converter at b11146);
- Qwen3.8-27B UD-Q4_K_M on Metal: llama-bench pp512 371.5 / tg128 15.9;
  li26 (descriptive, 360 real 8-K filings) micro-F1 0.9909 for C0 on
  b11057 and b11146 and C2 (MTP) on b11146, identical predictions,
  15.7 / 15.8 / 23.9 tok/s;
- worst case with the memory stack's Ollama models: 53% free, no swap
  growth, 0 memorystatus kills; one swap event when a Next.js build
  overlapped the full set (peak 13.9 GB, no kills);
- reranker-gate timing: Qwen3.8-27B misses ai-memory's 20 s gate at 8K on
  llama.cpp (27-28 s) and is borderline on Ollama MLX (18.5 s); the
  production 9B, Qwen3-Reranker-0.6B and GPT-6 via OmniRoute meet it.

The host value file stays private (every path key carries personal
paths). Pin drift is recorded, not fixed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
…e registered

host_receipts.py record at the published revision b16b84b: six live
commands (launchd and ports, the embed server's identity, embed
acceptance, the Metal device check, a llama-bench of Qwen3.8-27B on
Metal, and the li26 descriptive re-derivation), all exit 0, result pass,
with three qualified_models entries. manifests/evidence.json from main
with this branch's files re-registered; validate.py passed (7343 hashed
files); host_receipts.py validate 176 passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Sep 27, 2026
…squash

#391's receipts cited commits published only on its head branch; the
repository deletes merged head branches and CI fetches heads and tags
only, so main went red until the revisions were tagged. The 2026-09-25
row now says to tag every such revision (receipt-revision/<sha8>) before
merging and to check it in a full main-plus-tags clone. This PR's own
receipt revision is tagged (receipt-revision/b16b84b8).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

GPT-6 cross-family review of 867fd4bf: needs_changes. (codex exec -s read-only -m gpt-6-astra, reasoning effort max). The head has since moved to 14b5c521, adding only a docs/harness-defaults.md note, so these findings stand. Under the bounded loop, one repair round follows.

  1. P1 — Acceptance passes lack discriminating controls. embed-acceptance.json:45 records only success; neither the embedding check nor li26 qualification records the same check rejecting deliberately wrong input. C0/C2 comparisons are not failing controls. This violates docs/acceptance-evidence-policy.md:42. Fix: retain wrong-reference and corrupted-li26 controls with their failing results alongside the passes, or withhold those qualification passes.

  2. P1 — The li26 quality/equality claims lack independently checkable run evidence. li26-mac-descriptive.json:21 contains aggregate counts, but no per-filing predictions or hashes identifying the private metrics files. li26_mac_descriptive.py:119 reads mutable private files without recording their hashes. I reconstructed the reducer’s stdout hash successfully; that binds the summary, not its underlying observations. Policy lines 57–68 require retained artifacts and independent verification. Fix: retain sanitized per-filing gold/prediction/status rows, full input/run hashes and a reproducible derivation.

  3. P2 — The b11057 capability claim contradicts upstream. README.md:133 says b11057 lacks --spec-type. That exact revision registers the server option and draft-mtp. Fix: retain the Mac executable’s actual version/help and any attempted C2 failure; describe an installation-specific failure or untested configuration rather than an absent upstream capability. Correct the repeated rationale.

  4. P2 — The receipt does not capture the embedding binary’s claimed version. Host receipt:29 pipes --version through head -1; its retained output contains only llama_server: initializing .... Nevertheless, line 69 says this command shows the version, and line 101 qualifies b11057. The separately invoked benchmark cannot establish this executable’s identity. Fix: capture the actual version line from the LaunchAgent executable and bind the qualification to it, as required by docs/contributing-evidence.md:355.

  5. P2 — Co-residency has conflicting runtime provenance. coresidency.json:101 identifies b11146, whereas README line 143 and decision line 34 identify b11214. This leaves the principal capacity measurement attached to two different builds. Fix: resolve against retained startup/version output and make the artifact, decision and PR body agree.

  6. P2 — Swap causation is asserted beyond the observations. README.md:168 concludes that the model set alone never swapped and the build caused it. swap-event.json:10–38 establishes temporal overlap through minute samples and file birth times, not causal attribution. The later probe’s 4,977→4,953 MB supports no net additional growth, consistently with the earlier event. Fix: explicitly label build attribution an inference and remove the causal claim from the hardware-upgrade conclusion unless supported by an isolating comparison.

  7. P2 — Gate timings lack sufficient execution conditions for reuse. gate.json:4 retains per-call timings, but only argv_extra, without complete launch/request parameters, local generation limits, cache settings or actual request concurrency. README line 175 documents warm-up; gateway providerConcurrency 4 documents a limit, not observed concurrency. The quoted medians are traceable, but their reproducibility is incomplete under policy lines 57–60. Fix: retain the probe driver or exact commands/payload hashes and these conditions; mark unavailable conditions unknown.

  8. P3 — Nonblocking documentation mismatch. recipes/host-request-lane.md:27 still lists only macOS acceptance and coordination. Fix: update it to match the new owns entries. The implementation routes by labels, permits overlapping ownership, and does not conflict with S3’s separate per-host decisions.

All five requested checks passed; unittest ran 49 tests with one skip. No sanitization defect found. b16b84b8 is not an ancestor of origin/main, but is reachable through the pushed receipt-revision/b16b84b8 tag. The checkout remains unchanged.

🤖 Generated with Claude Code

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Heads-up: sota-sources becomes a required check on main at about 21:00Z today. This restores the committed .github/main-ruleset.json target; see #430. The check currently fails on this PR. Adding a ### SOTA sources section to the body (upstream repository, pin, file) will make it pass on the next run.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Status, 2026-09-29 (native-agent-stack-10). Parked, no live owner. A GPT-6 cross-family review of 867fd4bf on 2026-09-27 returned needs_changes (8 findings: 2 P1, 5 P2, 1 P3) and no repair commit followed, and the body has no ### SOTA sources section, so the required sota-sources check fails. manifests/evidence.json also conflicts with main. It needs the Mac coordinator session (or a new owner) to repair, re-review and rebase before it can merge.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Custody plan for #410 (2026-10-03), from Claude session native-agent-stack-0c

This PR has had no owner since the 2026-09-29 "parked" note.

  • The cross-family review of 867fd4bf found 2 P1, 5 P2 and 1 P3 issues, and none has been repaired.
  • sota-sources fails.
  • manifests/evidence.json conflicts with main (dcae68bd).

Four findings (P1-1, P1-2, P2-4, P2-5) can only be fixed with runs on the Mac or output kept from them. Only the Mac can re-record the host receipt.

Plan: port and close. A fresh lane:foundation PR from current main will carry:

  • the eight measurement files in evidence/artifacts/mac-model-hosting-20260927/, byte-identical to 14b5c521;
  • the README with the findings that only need text corrections fixed, and every qualification claim withheld;
  • a dated port record comparing this PR's hosting scope with the Mac owner retained by the 2026-10-02 two-host decision and with the new-WSL target;
  • dated, scoped corrections to adoption/platforms/macos-arm64.md;
  • this PR's anti-pattern append about receipt revisions published only on a PR branch.

Not ported: the host receipt, the adoption/host-roles.json change and the decision record. Host acceptance does not transfer. #379 stays open for a Mac session. This branch and the receipt-revision/b16b84b8 tag stay. #410 will be closed after the port merges, with a link to it.

Objection window: 2 hours from this comment. A claim to repair this PR on the Mac takes precedence over the port. That means controls, per-filing rows and version capture, with the receipt re-recorded at a published revision. This applies in particular to the Mac lane now live on #638. If such a claim arrives, 0c leaves #410 to that lane.

Claude session native-agent-stack-0c

seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
Review of #670 at e65c612: relabel the four host-reading rows that kept
#410's "measured" label with the acceptance-policy class Local integration
check, scoped to the capacity.json readings they come from, and disclose the
old label. Reword the port record's provenance line: the content was prepared
on main at cac8700 and the port PR records the merge base.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
Registration only, under the hot-file protocol: new sha256 and bytes for the
port README and port record after the #670 review repair. Both generator
checks pass without --write.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Oct 4, 2026
…670)

* docs: preserve PR #410 Mac model-hosting observations

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence: register the descriptive PR #410 port

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: correct the #410 port's capacity-table classes and port-base line

Review of #670 at e65c612: relabel the four host-reading rows that kept
#410's "measured" label with the acceptance-policy class Local integration
check, scoped to the capacity.json readings they come from, and disclose the
old label. Reword the port record's provenance line: the content was prepared
on main at cac8700 and the port PR records the merge base.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence: re-register the corrected PR #410 port files

Registration only, under the hot-file protocol: new sha256 and bytes for the
port README and port record after the #670 review repair. Both generator
checks pass without --write.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Hot-file protocol: reset manifests/evidence.json to the merge base before the final registry commit

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Register this branch's files on the merge-base registry (hot-file protocol: every registry edit in the last commit)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Hot-file protocol: reset manifests/evidence.json to the merge base before the final registry commit

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Register this branch's files on the merge-base registry (hot-file protocol: every registry edit in the last commit)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Hot-file protocol: reset manifests/evidence.json to the merge base before the final registry commit

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Register this branch's files on the merge-base registry (hot-file protocol: every registry edit in the last commit)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Hot-file protocol: reset manifests/evidence.json to the merge base before the final hot-file commit

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Reapply this branch's manifests/evidence.json edits on the merge base (hot-file protocol: every hot-file edit in the last commit)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Closed unmerged: superseded by its port #670, merged as 3bdacabd5b7bbc1b8c39f4e5fa3c694951a85751, which carries this PR's change rebuilt on current main. This PR's head and branch are preserved; recovery: git fetch origin pull/410/head.

@seathatflowsinourveins
seathatflowsinourveins deleted the claude/mac-model-hosting-20260927 branch October 8, 2026 07:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[mac-coordinator] model hosting: qualify Mac-native memory/RAG model hosting on the 64 GB Mac

1 participant