Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
17 commits
Select commit Hold shift + click to select a range
c3b5f24
llamacpp: GPU offload, sequence-copy fan-out, marginal readout
maximelefrancois86 Sep 23, 2026
c0fbc50
llamacpp: size the fan-out from the data, offload by default; GGUF re…
maximelefrancois86 Sep 23, 2026
d4d26fd
docs: say what the marginal readout is and is not shown to do; where …
maximelefrancois86 Sep 23, 2026
c550dd5
docs: the reproduce loop ran shared mode on a multi-state file; drop it
maximelefrancois86 Sep 23, 2026
01649b8
results: Torch BF16 on the same laptop GPU, with and without the fast…
maximelefrancois86 Sep 23, 2026
5f98fe0
results: the line the PR starts from — llama.cpp on CPU as merged, be…
maximelefrancois86 Sep 23, 2026
b7a1ffb
llamacpp: drop the marginal readout — on Qwen3-8B under this prompt i…
maximelefrancois86 Sep 23, 2026
5e546b5
Add semif-calibrate CLI and generic calibrated-threshold gate
cs-fisha Sep 25, 2026
2a5a2ef
Add option-order sensitivity measurement and opt-in stabilize-order
cs-fisha Sep 25, 2026
fc32e7e
Trim the shared state prefix to a real token boundary
pxtroniwnl Sep 26, 2026
d04a709
Cover the object-state boundary variant in the prefix regression
pxtroniwnl Sep 26, 2026
d3de88f
Add an SGLang server backend for all four scoring modes
rwang5203 Sep 28, 2026
fb12eba
Merge branch 'pr-41' into picks
pavan-rsa-ecom Oct 2, 2026
58de8cb
Merge branch 'pr-42' into picks
pavan-rsa-ecom Oct 2, 2026
b0344cc
Merge branch 'pr-34' into picks
pavan-rsa-ecom Oct 2, 2026
b781152
Merge branch 'pr-55' into picks
pavan-rsa-ecom Oct 2, 2026
b7bcb30
fix(order): sample permutations lazily above 8 options
pavan-rsa-ecom Oct 2, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 13 additions & 11 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,20 +54,20 @@ export HF_HOME=/path/to/large-drive/huggingface
pip install -e '.[test]'
```

**CPU only:** the llama.cpp backend scores the same prompts from a local GGUF
checkpoint with no CUDA device. Install `pip install -e '.[test,llamacpp]'`,
fetch a GGUF (for example `Qwen_Qwen3.5-4B-Q4_K_M.gguf` from
`bartowski/Qwen_Qwen3.5-4B-GGUF`), and add `--backend llamacpp --gguf
/path/to/model.gguf`; `--llama-threads` caps the CPU threads. Prompt
construction stays on the pinned reference tokenizer, so `prompt_sha256`
matches the Torch backend row for row; scores carry the GGUF checksum and are
conditional on the quantized weights. Direct and prefix-cached execution can
have small numerical differences from different llama.cpp evaluation paths;
compare decisions or probabilities with a tolerance rather than raw logits
bit for bit. One loaded backend owns one stateful scoring context. For a much
**llama.cpp / GGUF:** the llama.cpp backend scores the same prompts from a local GGUF
checkpoint, on CPU or with `--llama-gpu-layers` offloaded to a GPU. Install
`pip install -e '.[test,llamacpp]'`, download a GGUF of the pinned model (for example
`bartowski/Qwen_Qwen3.5-4B-GGUF`), and add `--backend llamacpp --gguf /path/to/model.gguf`.
Layers are offloaded by default when the wheel can, and shared mode fans a state's questions
out over copied sequences in as few batched decodes as the context allows — sized from the
rows themselves, which also works on hybrid Qwen3.5 memories. Prompt hashes match the
Torch backend; quantized option scores have small numerical differences. See
[docs/LLAMACPP.md](docs/LLAMACPP.md). For a much
slower full-precision Torch reference path, explicitly pass
`--device cpu --dtype float32` to the standard scorer command.

**SGLang server:** `--backend sglang` scores the same prompts through a running SGLang server that contains sgl-project/sglang#40826 (SGLang main from commit 174a5f37 of 2026-09-24 on, or a nightly from 0.5.21.dev20260925 on). The pinned reference tokenizer still builds every prompt, so `prompt_sha256` matches the Torch backend row for row, and the server scores those token ids on its own GPU. `--sglang-url` names the server (default `http://127.0.0.1:30000`). Scores can differ from Torch, and shared mode on SGLang does not guarantee a single prefill of the state. The [SGLang guide](docs/SGLANG.md) has the server launch and the refused server settings.

Run the owned examples:

```bash
Expand Down Expand Up @@ -190,7 +190,9 @@ Returned probabilities are conditional on the supplied options. Calibrate and va
- [Method](docs/METHOD.md) — frozen prompts, metrics, and timing scope
- [Reproduce](docs/REPRODUCE.md) — exact environment, pinned commands, perturbations, and verification
- [Apple Silicon](docs/APPLE_SILICON.md) — MPS and optional MLX backends
- [SGLang server backend](docs/SGLANG.md)
- [Calibration](docs/CALIBRATION.md) — fitted temperatures, out-of-fold evidence, and application
- [Option order](docs/OPTION_ORDER.md) — order-sensitivity measurement and opt-in `--stabilize-order`
- [EXL3 bridge](exl3-bridge/README.md) — quantized 27B runner and committed evidence
- [Interactive replay](demo/index.html)
- [Browser-only WebGPU demo](webgpu-demo/index.html) — no waitlist; use it today
Expand Down
25 changes: 25 additions & 0 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,19 @@ python benchmarks/build_perturbations.py \

`docs/REPRODUCE.md` gives the complete command for rebuilding `results/raw/perturbation-comparison.json` from the committed row-level predictions. The regenerated report is byte-identical to the committed report.

Recompute the option-order sensitivity digest (flips, total variation, position bias) from those same committed files without loading a model:

```bash
python benchmarks/evaluate_option_order.py \
--gold benchmarks/data/authored144.jsonl \
--perturbations benchmarks/data/perturbations108.jsonl \
--base-predictions results/raw/predictions/direct-authored144.jsonl \
--perturbation-predictions results/raw/predictions/direct-perturbations108.jsonl \
--output /tmp/option-order-direct.json
```

See [OPTION_ORDER.md](../docs/OPTION_ORDER.md) for the opt-in `--stabilize-order K` scorer flag and its `K×` cost. Do not mix stabilize-order outputs into frozen quality tables.

## Full 37×21 systems benchmark

```bash
Expand All @@ -48,6 +61,18 @@ CUDA_VISIBLE_DEVICES=0 python benchmarks/shape777_reranker.py \
--output shape777-reranker-run.json
```

The same fixture on a local GGUF through llama.cpp, layers offloaded by default when the wheel
can, branches sized per state:

```bash
python benchmarks/shape777.py --backend llamacpp \
--model Qwen/Qwen3.5-4B \
--revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
--gguf /path/to/Qwen_Qwen3.5-4B-Q4_K_M.gguf \
--input benchmarks/data/shape777.jsonl \
--output shape777-llamacpp-run.json
```

The 6.7 MB fixture is project-authored and has SHA-256 `8dcf414b12fc2684e3c4ca5f3ebfd3f525f5346fec4a9bc67eb65138101f55f1`. Both runners write aggregate timings and row-level predictions.

## Quality evidence
Expand Down
Loading