Skip to content

Add Apple Metal (MPS) and CPU support with auto device selection - #6

Open
james-see wants to merge 3 commits into
TheoLeeCJ:masterfrom
james-see:master
Open

james-see wants to merge 3 commits into
TheoLeeCJ:masterfrom
james-see:master

Conversation

@james-see

Copy link
Copy Markdown
Contributor

Summary

Scorer and benchmark runners now work on Apple Silicon (MPS) and CPU in addition to CUDA, with automatic device selection (CUDA > MPS > CPU).

Changes

  • core.py: new `resolve_device()`, `resolve_dtype()`/`dtype_candidates()` (bf16 preferred, fp16/fp32 fallback off CUDA), `synchronize()`, `describe_hardware()`, portable peak-memory helpers. `load_causal_model()` accepts optional `device`/`dtype`; CUDA path stays strict (exactly-one-GPU, bf16-only) so published results remain comparable; actual dtype/device recorded in metadata.
  • direct/reranker/shared/serial: `torch.cuda.synchronize` branches replaced with portable `synchronize(device)`.
  • CLI + shape777 / shape777_reranker / decision_vs_generation: new `--device`/`--dtype` flags (default `auto`); portable hardware/memory reporting (`peak_cuda_bytes` kept for compat, plus `peak_memory_bytes` + `device`).
  • Docs: AGENTS.md one-accelerator policy; benchmarks/README Apple Silicon usage; METHOD.md MPS/CPU non-comparability note.
  • tests/test_device.py: 14 tests for resolution priority, dtype policy, sync safety, revision validation.

Validation

  • `pytest tests/ -q`: 26 passed (incl. live MPS smoke on ARM64 Mac)
  • `(cd results/raw && sha256sum -c SHA256SUMS)`: 15 OK
  • `python benchmarks/verify_published.py`: 69 claims, status ok
  • No headline claims, phase1-summary.json, or committed raw outputs touched."

@james-see

Copy link
Copy Markdown
Contributor Author

@TheoLeeCJ this should be gtg now no conflicts. need to support MPS / mac metal vs. just CUDA

@james-see

Copy link
Copy Markdown
Contributor Author

@TheoLeeCJ any movement on this?

TheoLeeCJ added a commit that referenced this pull request Sep 23, 2026
Port the focused CPU device-selection work from PR #6 onto the current
backends while keeping automatic selection accelerator-only and the
reranker CUDA-only.

Co-authored-by: James Campbell <616585+james-see@users.noreply.github.com>
@TheoLeeCJ

Copy link
Copy Markdown
Owner

Sorry for the wait. Other PRs now cover Apple Silicon and llama.cpp CPU support, but the Torch CPU and device-selection work here remained useful. I ported that focused portion onto current master in 23cf1f3 and preserved attribution to your contribution. Thank you for the work!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants