Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 9 additions & 11 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,17 +54,15 @@ export HF_HOME=/path/to/large-drive/huggingface
pip install -e '.[test]'
```

**CPU only:** the llama.cpp backend scores the same prompts from a local GGUF
checkpoint with no CUDA device. Install `pip install -e '.[test,llamacpp]'`,
fetch a GGUF (for example `Qwen_Qwen3.5-4B-Q4_K_M.gguf` from
`bartowski/Qwen_Qwen3.5-4B-GGUF`), and add `--backend llamacpp --gguf
/path/to/model.gguf`; `--llama-threads` caps the CPU threads. Prompt
construction stays on the pinned reference tokenizer, so `prompt_sha256`
matches the Torch backend row for row; scores carry the GGUF checksum and are
conditional on the quantized weights. Direct and prefix-cached execution can
have small numerical differences from different llama.cpp evaluation paths;
compare decisions or probabilities with a tolerance rather than raw logits
bit for bit. One loaded backend owns one stateful scoring context.
**llama.cpp / GGUF:** the llama.cpp backend scores the same prompts from a local GGUF
checkpoint, on CPU or with `--llama-gpu-layers` offloaded to a GPU. Install
`pip install -e '.[test,llamacpp]'`, download a GGUF of the pinned model (for example
`bartowski/Qwen_Qwen3.5-4B-GGUF`), and add `--backend llamacpp --gguf /path/to/model.gguf`.
Layers are offloaded by default when the wheel can, and shared mode fans a state's questions
out over copied sequences in as few batched decodes as the context allows — sized from the
rows themselves, which also works on hybrid Qwen3.5 memories. Prompt hashes match the
Torch backend; quantized option scores have small numerical differences. See
[docs/LLAMACPP.md](docs/LLAMACPP.md).

Run the owned examples:

Expand Down
12 changes: 12 additions & 0 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,18 @@ CUDA_VISIBLE_DEVICES=0 python benchmarks/shape777_reranker.py \
--output shape777-reranker-run.json
```

The same fixture on a local GGUF through llama.cpp, layers offloaded by default when the wheel
can, branches sized per state:

```bash
python benchmarks/shape777.py --backend llamacpp \
--model Qwen/Qwen3.5-4B \
--revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
--gguf /path/to/Qwen_Qwen3.5-4B-Q4_K_M.gguf \
--input benchmarks/data/shape777.jsonl \
--output shape777-llamacpp-run.json
```

The 6.7 MB fixture is project-authored and has SHA-256 `8dcf414b12fc2684e3c4ca5f3ebfd3f525f5346fec4a9bc67eb65138101f55f1`. Both runners write aggregate timings and row-level predictions.

## Quality evidence
Expand Down
63 changes: 55 additions & 8 deletions benchmarks/shape777.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,9 +11,18 @@
from pathlib import Path

from semif_phase1.core import load_causal_model
from semif_phase1.direct import score
from semif_phase1.serial import SerialPrefixScorer
from semif_phase1.shared import score_shared
from semif_phase1.direct import score as torch_score
from semif_phase1.serial import SerialPrefixScorer as TorchSerialPrefixScorer
from semif_phase1.shared import score_shared as torch_score_shared


def _gpu_name() -> str | None:
"""The NVIDIA device name without importing a CUDA-enabled torch."""
for info in sorted(Path("/proc/driver/nvidia/gpus").glob("*/information")):
for line in info.read_text().splitlines():
if line.startswith("Model:"):
return line.split(":", 1)[1].strip()
return None


def main() -> None:
Expand All @@ -23,17 +32,53 @@ def main() -> None:
parser.add_argument("--input", type=Path, required=True)
parser.add_argument("--output", type=Path, required=True)
parser.add_argument("--max-tokens", type=int, default=4096)
parser.add_argument("--backend", choices=("torch", "llamacpp"), default="torch")
parser.add_argument("--gguf", type=Path, help="Local GGUF checkpoint for --backend llamacpp")
parser.add_argument("--llama-threads", type=int, help="CPU threads for --backend llamacpp")
parser.add_argument("--llama-gpu-layers", default="auto",
help="GPU layers for --backend llamacpp: 'auto' (library default), 0 (CPU) or N")
parser.add_argument("--llama-parallel", default="auto",
help="llama.cpp shared-mode branching: 'auto' (sized per state), N, or 1 (state restore)")
args = parser.parse_args()
if args.output.exists():
parser.error("Output must be new")
if args.backend == "llamacpp":
if args.gguf is None or not args.gguf.is_file():
parser.error("--backend llamacpp requires --gguf pointing at an existing GGUF file")
for name in ("llama_gpu_layers", "llama_parallel"):
value = getattr(args, name)
if value != "auto":
try:
setattr(args, name, int(value))
except ValueError:
parser.error(f"--{name.replace('_', '-')} must be 'auto' or an integer")
elif args.gguf is not None or args.llama_threads is not None or args.llama_gpu_layers != "auto" or args.llama_parallel != "auto":
parser.error("llama.cpp options require --backend llamacpp")
rows = [json.loads(line) for line in args.input.read_text().splitlines() if line.strip()]
groups = defaultdict(list)
for row in rows:
groups[row["group_id"]].append(row)
if len(rows) != 777 or len(groups) != 37 or any(len(group) != 21 for group in groups.values()):
parser.error("Expected the committed 37-state x 21-question fixture")
model, tokenizer, metadata = load_causal_model(args.model, args.revision, "cuda")
import torch
if args.backend == "llamacpp":
from semif_phase1 import llamacpp_backend

model, tokenizer, metadata = llamacpp_backend.load_model(
args.model, args.revision, args.gguf, threads=args.llama_threads,
context_tokens=args.max_tokens, gpu_layers=args.llama_gpu_layers,
sequences=args.llama_parallel)
score = llamacpp_backend.score
SerialPrefixScorer = llamacpp_backend.SerialPrefixScorer
score_shared = llamacpp_backend.score_shared
cuda = None
offloaded = metadata["n_gpu_layers"] != 0 and metadata["gpu_offload_supported"]
hardware = (_gpu_name() or "unknown GPU") if offloaded else "CPU"
else:
model, tokenizer, metadata = load_causal_model(args.model, args.revision, "cuda")
import torch as cuda

score, SerialPrefixScorer, score_shared = torch_score, TorchSerialPrefixScorer, torch_score_shared
hardware = cuda.cuda.get_device_name(0)

first = next(iter(groups.values()))
score(model, tokenizer, first[0], metadata, args.max_tokens)
Expand All @@ -45,13 +90,15 @@ def main() -> None:
"version": "shape777-published-v1",
"input_sha256": hashlib.sha256(args.input.read_bytes()).hexdigest(),
"model": metadata,
"hardware": torch.cuda.get_device_name(0),
"hardware": hardware,
"backend": args.backend,
"timing_scope": "Warm model; includes prompt construction, tokenization, transfers, forward passes and CPU readout.",
"results": [],
}
predictions = {}
for mode in ("fresh", "serial_prefix", "parallel_shared"):
torch.cuda.reset_peak_memory_stats()
if cuda is not None:
cuda.cuda.reset_peak_memory_stats()
started = time.perf_counter()
values, state_times = [], []
for group in groups.values():
Expand All @@ -73,7 +120,7 @@ def main() -> None:
"wall_seconds": elapsed,
"decisions_per_second": len(values) / elapsed,
"state_p50_seconds": statistics.median(state_times),
"peak_cuda_bytes": torch.cuda.max_memory_allocated(),
"peak_cuda_bytes": cuda.cuda.max_memory_allocated() if cuda is not None else None,
}
)
reference = {row["id"]: row for row in predictions["fresh"]}
Expand Down
227 changes: 227 additions & 0 deletions docs/LLAMACPP.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,227 @@
# llama.cpp / GGUF

The llama.cpp backend runs SemIf's direct, serial-prefix, and shared decision
modes over a quantized GGUF checkpoint, on CPU or with layers offloaded to a
GPU. Prompts, answer-slot checks and `prompt_sha256` come from the reference
transformers tokenizer, exactly as in the Torch backend; llama.cpp only executes
the forward pass. Every prompt is re-tokenized through the GGUF vocabulary and
must match the reference encoding before it is scored. No answer token is
generated.

## Install and score

```bash
python3 -m venv .venv
source .venv/bin/activate
pip install -e '.[test,llamacpp]' # CPU wheel of llama-cpp-python

semif-score --backend llamacpp --mode direct \
--model Qwen/Qwen3.5-4B \
--revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
--gguf /path/to/Qwen_Qwen3.5-4B-Q4_K_M.gguf \
--input examples/decisions.jsonl \
--output results-llamacpp-direct.jsonl
```

The pinned tokenizer is downloaded into the Hugging Face cache on first use
(a few megabytes; the model weights come from the GGUF file). The GGUF used for
the results below is `bartowski/Qwen_Qwen3.5-4B-GGUF` at revision
`4168f45a16a1290d65a4ec0fa312ae917a4c15d6`, quantization Q4_K_M, 3 013 027 808
bytes, SHA-256 `13c16f426047e2de38cd075bdade4a7bcbc8c774384876f677740cda65f8a983`.
Each prediction records the file's size and checksum.

### GPU offload

`--llama-gpu-layers` defaults to `auto`, which leaves llama.cpp's own default —
`-1`, every layer, in current builds — and records what the library did. `0`
forces CPU, `N` offloads `N` layers. Offload needs a llama-cpp-python wheel
built with a GPU backend; the PyPI wheel is CPU-only and ignores the setting
(`gpu_offload_supported` in the metadata says so). For CUDA:

```bash
pip install --force-reinstall --no-deps "llama-cpp-python==0.3.35" \
--extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu124
pip install nvidia-cuda-runtime-cu12 nvidia-cublas-cu12 nvidia-cuda-nvrtc-cu12
export LD_LIBRARY_PATH=$(ls -d .venv/lib/python3.*/site-packages/nvidia/*/lib | tr '\n' ':')
```

The metadata records `n_gpu_layers` and whether the loaded library supports
offload at all, so a CPU wheel silently ignoring the flag is visible in the
output.

## Two ways to branch from one state

Sequence 0 of the llama.cpp context holds the prefilled state. Shared mode
scores every question of that state from a branch of it, and there are two
ways to make a branch:

| `--llama-parallel` | branching | decodes per state |
|---|---|---|
| `auto` (default) | **sequence copy, sized per state** — as many of the state's questions per decode as the context holds: `prefix + Σ suffixes ≤ n_ctx`, at most 32 branches | usually one |
| `N ≥ 2` | sequence copy with `n_seq_max = N`, chunks of at most `N − 1` | one per `N − 1` questions |
| `1` | **state restore** — `llama_state_seq_get_data` after the prefill, `set_data` before each question | one per question |

State restore is what the backend has always done. Sequence copy is the
llama.cpp equivalent of the Torch backend's `native-state-prefix-parallel-v1`:
the prefix is evaluated once, copied without serialization, and the branch
suffixes share a single decode. The branches are removed whole afterwards
(whole-sequence removal never fails), so sequence 0 is left exactly as
prefilled and the next chunk of questions needs no restore at all.

Why `auto` needs no number from you: everything that bounds a fan-out is known
before the first decode. `encode_verified` has already tokenized each row, so
the prefix length and every suffix length are in hand, and the context budget
is fixed at load. The branch count is therefore a consequence of the data — on
the `shape777` fixture (states of ~1 800 tokens, suffixes of ~90) all 21
questions of a state fit in one decode. `branches_per_decode` in the shared
timing records what was actually done.

The context is sized **once**, for the longest single prompt (`max_tokens + 64`),
not multiplied by the number of sequences. That relies on `kv_unified`: measured
on this build, with 16 sequences and `n_ctx = 4096`, each sequence sees 4 096
cells under the unified buffer and 256 without it, and copied branches share the
prefix's cells rather than duplicating them. A library without `kv_unified` falls
back to the per-sequence allocation.

### Hybrid models

Qwen3.5 is a hybrid architecture: one full-attention layer in four, the rest
Gated DeltaNet with a recurrent state. In llama.cpp that state lives in
`llama_memory_recurrent`, which keeps only the state after the last token and
therefore **cannot be partially erased** — `llama_memory_seq_rm` refuses any
range that includes a sequence's last position (the source says: "models like
Mamba or RWKV can't have a state partially erased at the end of the sequence").
Truncating a branch back to the prefix is thus impossible on these models, and
that is a property of the architecture, not of a build or a binding.

What hybrid memories do support is whole-sequence removal, state save/restore,
and **sequence copies** — `llama_memory_recurrent::seq_cp` is implemented — which
is why both branching modes above work on Qwen3.5. Pure-attention models
additionally allow tail truncation; the backend does not rely on it.

`n_seq_max` is a context parameter (`llama_context_params.n_seq_max`), not a
build option; the same wheel serves both modes.

## Validation

`tests/test_llamacpp.py::test_real_gguf_scores_direct_serial_and_shared` loads the
real GGUF twice — with one sequence and with three — and checks that direct,
serial, restore-shared and copy-shared readouts agree on the decisions and
that copy-shared option logits match restore-shared ones within 0.5. Set `SEMIF_LLAMACPP_GGUF` to run
it, `SEMIF_LLAMACPP_GPU_LAYERS` to offload. It passes on:

- the PyPI CPU wheel of `llama-cpp-python` 0.3.35, CPU only;
- the cu124 wheel of the same version with every layer offloaded to an RTX 3080
Laptop GPU.

**Known issue, not in this backend:** the cu124 wheel's *CPU* code path raises
`Illegal instruction` on an Intel Xeon W-11955M (AVX-512 without `avx512_bf16`
or AMX) for any decode, single- or multi-sequence, hybrid or pure-attention
model. With that wheel, offload the layers; for CPU scoring, use the PyPI wheel.

## Results

RTX 3080 Laptop GPU (16 GB), Q4_K_M, all 33 layers offloaded, `llama-cpp-python` 0.3.35 cu124.

### Quality — `authored144`, direct mode

| backend | GPU | mean family balanced accuracy | ECE, own-T out-of-fold | fitted T |
|---|---|---:|---:|---:|
| Torch BF16 (committed) | RTX 3090 | 0.813 | 0.038 | 1.23 |
| Torch BF16, same laptop | RTX 3080 Laptop | 0.813 | 0.050 [0.038, 0.114] | 1.25 |
| **llama.cpp GGUF Q4_K_M** | RTX 3080 Laptop | **0.796** | **0.063** [0.038, 0.123] | 1.26 |

Coverage 144/144. Median `allowed_token_mass` 0.9997, minimum 0.975, no row
below 0.9 — the quantized model answers with a letter as reliably as the BF16
one. Median forward time 64 ms per decision at a median 147 input tokens.
Row-level predictions are in
`results/raw/predictions/llamacpp-gguf-cuda-direct-authored144.jsonl`, the
report in `results/raw/llamacpp-gguf-cuda-authored144.json`, the calibration
in `results/raw/calibration/llamacpp-gguf-cuda-authored144.json`. The 4-bit
checkpoint gives up 1.7 points of balanced accuracy and is somewhat less
well-calibrated than BF16; its fitted temperature is nearly the same.

### Systems — `shape777`, 37 states × 21 questions

| backend, GPU | fresh | serial prefix | parallel shared | flips vs fresh |
|---|---:|---:|---:|---:|
| Torch BF16, RTX 3090 (committed) | 2.33 | 10.75 | 20.03 | 6 / 777 |
| Torch BF16, RTX 3080 Laptop, `flash-linear-attention` | 1.51 | 10.19 | 14.14 | 3 / 777 |
| Torch BF16, RTX 3080 Laptop, reference PyTorch kernels | 1.10 | 7.78 | 10.24 | 9 / 777 |
| **llama.cpp Q4_K_M, RTX 3080 Laptop**, `--llama-parallel 8` | **1.40** | **9.21** | **10.88** | 16 / 777 |
| llama.cpp Q4_K_M, RTX 3080 Laptop, `auto` | | | 10.51 | 18 / 777 |

Decisions per second; state p50 for the GGUF: 14.97 s fresh, 2.28 s serial,
1.93 s parallel.

What the table does not show is the line this PR starts from: the backend as
merged in #18 runs on CPU only. Measured on the first 5 states of the fixture
(105 decisions, backend functions timed directly because `shape777.py` takes
only the full fixture; `results/raw/shape777-subset5-llamacpp-before-after.json`):

| llama.cpp Q4_K_M, same laptop, 5 states | fresh | shared |
|---|---:|---:|
| before: CPU (16 threads), state restore — `#18` as merged | 0.05 | 0.49 |
| after: GPU, state restore (`--llama-parallel 1`) | 1.45 | 9.86 |
| after: GPU, fan-out (`--llama-parallel auto`) | 1.43 | 11.18 |

The GPU offload is the gain — ×30 fresh, ×20 shared. The fan-out adds 13 % on
this subset and 18 % on the full fixture. The GPU rows of the subset agree with
the full-fixture rows above (9.21 / 10.88), so the subset is representative. The two Torch rows on the laptop differ only by the Gated
DeltaNet kernels: `pip install -e .` leaves `transformers` on its reference
PyTorch implementation (it says so at load time); with `flash-linear-attention`
installed the delta rule runs in Triton. `causal_conv1d` needs `nvcc` and was
not installed. Neither kernel touches llama.cpp, which has its own ggml
operators for these layers. On the same GPU, BF16 with the fast kernel is 8 %,
11 % and 30 % faster than the 4-bit GGUF in the three modes; the GGUF trades
that for 3 GB of weights instead of 8.8 and an 8.8 GB CUDA peak instead of 11.8.

Two things the numbers say. Fan-out beats state restore by removing the
serialization round-trip, not by batching: one decode per state (`auto`,
21 branches) is no faster than three (8 branches), because llama.cpp splits a
batch into micro-batches of `n_ubatch` tokens either way and the recurrent
memory decodes them with `split_equal`. And the gap to the Torch parallel figure
is not the fan-out's: the Torch backend runs one dense BF16 forward over padded
suffixes, which this 4-bit hybrid path cannot match on a laptop GPU. Reports:
`results/raw/shape777-llamacpp-gguf-cuda.json` (8 branches),
`results/raw/shape777-llamacpp-gguf-cuda-auto.json`,
`results/raw/shape777-torch-bf16-rtx3080-laptop.json` and
`…-reference-kernels.json`, each with its `.predictions.jsonl`; the laptop
BF16 quality run is `results/raw/torch-bf16-rtx3080-laptop-authored144.json`
with its calibration report and predictions. The 16–19 argmax flips out of 777 between decode paths
are the quantized model's noise floor — two sequential reads of one prompt
already differ on about 3 % of decisions when only the micro-batch size
changes — not a property of the fan-out.


## JevBench

[JevBench](https://github.com/fstandhartinger/jevbench) runs SemIf through its `semif_direct`
adapter on the Torch backend (BF16, CUDA). This backend is what would let that adapter run the
pinned GGUF on a laptop GPU with shared states fanned out; it has not been run on the full public
set yet.

## Reproduce

```bash
# Quality: the owned labeled workload, direct and serial (shared needs one state per file)
for MODE in direct serial; do
semif-score --backend llamacpp --mode $MODE \
--model Qwen/Qwen3.5-4B --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
--gguf /path/to/Qwen_Qwen3.5-4B-Q4_K_M.gguf \
--input benchmarks/data/authored144.jsonl \
--output results/raw/predictions/llamacpp-gguf-cuda-$MODE-authored144.jsonl
done
python benchmarks/evaluate.py --gold benchmarks/data/authored144.jsonl \
--predictions results/raw/predictions/llamacpp-gguf-cuda-direct-authored144.jsonl

# Systems: the 37x21 fixture, fresh / serial-prefix / parallel-shared
python benchmarks/shape777.py --backend llamacpp \
--model Qwen/Qwen3.5-4B --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
--gguf /path/to/Qwen_Qwen3.5-4B-Q4_K_M.gguf \
--input benchmarks/data/shape777.jsonl \
--output results/raw/shape777-llamacpp-gguf-cuda.json
```

Shared mode requires every row of the input to carry the same exact state;
`shape777.py` groups its fixture by `group_id` and does that for you.
Loading