Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 12 additions & 11 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,17 +54,15 @@ export HF_HOME=/path/to/large-drive/huggingface
pip install -e '.[test]'
```

**CPU only:** the llama.cpp backend scores the same prompts from a local GGUF
checkpoint with no CUDA device. Install `pip install -e '.[test,llamacpp]'`,
fetch a GGUF (for example `Qwen_Qwen3.5-4B-Q4_K_M.gguf` from
`bartowski/Qwen_Qwen3.5-4B-GGUF`), and add `--backend llamacpp --gguf
/path/to/model.gguf`; `--llama-threads` caps the CPU threads. Prompt
construction stays on the pinned reference tokenizer, so `prompt_sha256`
matches the Torch backend row for row; scores carry the GGUF checksum and are
conditional on the quantized weights. Direct and prefix-cached execution can
have small numerical differences from different llama.cpp evaluation paths;
compare decisions or probabilities with a tolerance rather than raw logits
bit for bit. One loaded backend owns one stateful scoring context.
**llama.cpp / GGUF:** the llama.cpp backend scores the same prompts from a local GGUF
checkpoint, on CPU or with `--llama-gpu-layers` offloaded to a GPU. Install
`pip install -e '.[test,llamacpp]'`, download a GGUF of the pinned model (for example
`bartowski/Qwen_Qwen3.5-4B-GGUF`), and add `--backend llamacpp --gguf /path/to/model.gguf`.
Layers are offloaded by default when the wheel can, and shared mode fans a state's questions
out over copied sequences in as few batched decodes as the context allows — sized from the
rows themselves, which also works on hybrid Qwen3.5 memories. Prompt hashes match the
Torch backend; quantized option scores have small numerical differences. See
[docs/LLAMACPP.md](docs/LLAMACPP.md).

Run the owned examples:

Expand Down Expand Up @@ -182,13 +180,16 @@ Calibration does not change the selected option. The clear improvement is on WAN
Returned probabilities are conditional on the supplied options. Calibrate and validate them on the workload where they will make decisions.
`state` may also be a nonempty JSON object or array. Direct modes preserve it as structured JSON; reranker mode renders it as document text.

The words around the input — the system instruction and the JSON key names — come from a prompt: `--prompt en` (default, the published wording), `--prompt fr`, or a JSON file. The payload shape and the answer letters never change; `prompt_version` in every result names the wording. See [docs/PROMPTS.md](docs/PROMPTS.md).

## Documentation

- [Results](docs/RESULTS.md) — quality, speed, perturbations, and claim boundaries
- [Method](docs/METHOD.md) — frozen prompts, metrics, and timing scope
- [Reproduce](docs/REPRODUCE.md) — exact environment, pinned commands, perturbations, and verification
- [Apple Silicon](docs/APPLE_SILICON.md) — MPS and optional MLX backends
- [Calibration](docs/CALIBRATION.md) — fitted temperatures, out-of-fold evidence, and application
- [Prompts](docs/PROMPTS.md) — the pluggable prompt: built-in wordings, prompt files, what never changes
- [EXL3 bridge](exl3-bridge/README.md) — quantized 27B runner and committed evidence
- [Interactive replay](demo/index.html)
- [Browser-only WebGPU demo](webgpu-demo/index.html) — no waitlist; use it today
Expand Down
12 changes: 12 additions & 0 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,18 @@ CUDA_VISIBLE_DEVICES=0 python benchmarks/shape777_reranker.py \
--output shape777-reranker-run.json
```

The same fixture on a local GGUF through llama.cpp, layers offloaded by default when the wheel
can, branches sized per state:

```bash
python benchmarks/shape777.py --backend llamacpp \
--model Qwen/Qwen3.5-4B \
--revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
--gguf /path/to/Qwen_Qwen3.5-4B-Q4_K_M.gguf \
--input benchmarks/data/shape777.jsonl \
--output shape777-llamacpp-run.json
```

The 6.7 MB fixture is project-authored and has SHA-256 `8dcf414b12fc2684e3c4ca5f3ebfd3f525f5346fec4a9bc67eb65138101f55f1`. Both runners write aggregate timings and row-level predictions.

## Quality evidence
Expand Down
63 changes: 55 additions & 8 deletions benchmarks/shape777.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,9 +11,18 @@
from pathlib import Path

from semif_phase1.core import load_causal_model
from semif_phase1.direct import score
from semif_phase1.serial import SerialPrefixScorer
from semif_phase1.shared import score_shared
from semif_phase1.direct import score as torch_score
from semif_phase1.serial import SerialPrefixScorer as TorchSerialPrefixScorer
from semif_phase1.shared import score_shared as torch_score_shared


def _gpu_name() -> str | None:
"""The NVIDIA device name without importing a CUDA-enabled torch."""
for info in sorted(Path("/proc/driver/nvidia/gpus").glob("*/information")):
for line in info.read_text().splitlines():
if line.startswith("Model:"):
return line.split(":", 1)[1].strip()
return None


def main() -> None:
Expand All @@ -23,17 +32,53 @@ def main() -> None:
parser.add_argument("--input", type=Path, required=True)
parser.add_argument("--output", type=Path, required=True)
parser.add_argument("--max-tokens", type=int, default=4096)
parser.add_argument("--backend", choices=("torch", "llamacpp"), default="torch")
parser.add_argument("--gguf", type=Path, help="Local GGUF checkpoint for --backend llamacpp")
parser.add_argument("--llama-threads", type=int, help="CPU threads for --backend llamacpp")
parser.add_argument("--llama-gpu-layers", default="auto",
help="GPU layers for --backend llamacpp: 'auto' (library default), 0 (CPU) or N")
parser.add_argument("--llama-parallel", default="auto",
help="llama.cpp shared-mode branching: 'auto' (sized per state), N, or 1 (state restore)")
args = parser.parse_args()
if args.output.exists():
parser.error("Output must be new")
if args.backend == "llamacpp":
if args.gguf is None or not args.gguf.is_file():
parser.error("--backend llamacpp requires --gguf pointing at an existing GGUF file")
for name in ("llama_gpu_layers", "llama_parallel"):
value = getattr(args, name)
if value != "auto":
try:
setattr(args, name, int(value))
except ValueError:
parser.error(f"--{name.replace('_', '-')} must be 'auto' or an integer")
elif args.gguf is not None or args.llama_threads is not None or args.llama_gpu_layers != "auto" or args.llama_parallel != "auto":
parser.error("llama.cpp options require --backend llamacpp")
rows = [json.loads(line) for line in args.input.read_text().splitlines() if line.strip()]
groups = defaultdict(list)
for row in rows:
groups[row["group_id"]].append(row)
if len(rows) != 777 or len(groups) != 37 or any(len(group) != 21 for group in groups.values()):
parser.error("Expected the committed 37-state x 21-question fixture")
model, tokenizer, metadata = load_causal_model(args.model, args.revision, "cuda")
import torch
if args.backend == "llamacpp":
from semif_phase1 import llamacpp_backend

model, tokenizer, metadata = llamacpp_backend.load_model(
args.model, args.revision, args.gguf, threads=args.llama_threads,
context_tokens=args.max_tokens, gpu_layers=args.llama_gpu_layers,
sequences=args.llama_parallel)
score = llamacpp_backend.score
SerialPrefixScorer = llamacpp_backend.SerialPrefixScorer
score_shared = llamacpp_backend.score_shared
cuda = None
offloaded = metadata["n_gpu_layers"] != 0 and metadata["gpu_offload_supported"]
hardware = (_gpu_name() or "unknown GPU") if offloaded else "CPU"
else:
model, tokenizer, metadata = load_causal_model(args.model, args.revision, "cuda")
import torch as cuda

score, SerialPrefixScorer, score_shared = torch_score, TorchSerialPrefixScorer, torch_score_shared
hardware = cuda.cuda.get_device_name(0)

first = next(iter(groups.values()))
score(model, tokenizer, first[0], metadata, args.max_tokens)
Expand All @@ -45,13 +90,15 @@ def main() -> None:
"version": "shape777-published-v1",
"input_sha256": hashlib.sha256(args.input.read_bytes()).hexdigest(),
"model": metadata,
"hardware": torch.cuda.get_device_name(0),
"hardware": hardware,
"backend": args.backend,
"timing_scope": "Warm model; includes prompt construction, tokenization, transfers, forward passes and CPU readout.",
"results": [],
}
predictions = {}
for mode in ("fresh", "serial_prefix", "parallel_shared"):
torch.cuda.reset_peak_memory_stats()
if cuda is not None:
cuda.cuda.reset_peak_memory_stats()
started = time.perf_counter()
values, state_times = [], []
for group in groups.values():
Expand All @@ -73,7 +120,7 @@ def main() -> None:
"wall_seconds": elapsed,
"decisions_per_second": len(values) / elapsed,
"state_p50_seconds": statistics.median(state_times),
"peak_cuda_bytes": torch.cuda.max_memory_allocated(),
"peak_cuda_bytes": cuda.cuda.max_memory_allocated() if cuda is not None else None,
}
)
reference = {row["id"]: row for row in predictions["fresh"]}
Expand Down
Loading