Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

- Run commands from the repository root in an isolated environment installed with `pip install -e '.[test]'`.
- Validate changes with `pytest -q`, `(cd results/raw && sha256sum -c SHA256SUMS)`, and `python benchmarks/verify_published.py`.
- Benchmark outputs are create-only. Use a new output path and expose exactly one CUDA GPU per scorer process.
- Benchmark outputs are create-only. Use a new output path and expose exactly one accelerator per scorer process (one CUDA GPU via CUDA_VISIBLE_DEVICES on NVIDIA hosts; MPS or CPU elsewhere).
- Do not change headline claims or `results/phase1-summary.json` without committing the supporting row-level evidence, regenerating the relevant raw report, updating `results/raw/SHA256SUMS`, and updating the method/results text.
- Preserve exact model and source revisions. Use `benchmarks/fetch_sources.py` only for its listed redistributable inputs; do not commit model weights, caches, or third-party raw records.
- `webgpu-demo/` is static and has no build step. Preserve `_headers`, runtime version pins, browser-only inference, and the explicit probability limitations.
15 changes: 10 additions & 5 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,10 +2,12 @@

The repository includes the exact owned speed fixture, benchmark runners, row-level model outputs, and source-selection IDs. Model weights and third-party records without a redistribution grant remain upstream.

Device selection is automatic: CUDA when exactly one GPU is visible, otherwise Apple Metal (MPS) when available, otherwise CPU. On NVIDIA hosts prefix GPU commands with `CUDA_VISIBLE_DEVICES=0`; on Apple Silicon run the same commands without it. Every runner also accepts `--device auto|cuda|mps|cpu` and `--dtype auto|bfloat16|float16|float32` (auto prefers bfloat16 with float16/float32 fallback off CUDA). Published timings used one RTX 3090 with BF16; MPS/CPU runs are functionally equivalent but not timing-comparable, and MPS may record a fallback dtype in `model.dtype`.

## Compact generation comparison

```bash
CUDA_VISIBLE_DEVICES=0 python benchmarks/decision_vs_generation.py \
python benchmarks/decision_vs_generation.py \
--model Qwen/Qwen3.5-4B \
--revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
--input benchmarks/data/shape777.jsonl \
Expand All @@ -30,17 +32,18 @@ python benchmarks/build_perturbations.py \
## Full 37×21 systems benchmark

```bash
CUDA_VISIBLE_DEVICES=0 python benchmarks/shape777.py \
python benchmarks/shape777.py \
--model Qwen/Qwen3.5-4B \
--revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
--input benchmarks/data/shape777.jsonl \
--output shape777-run.json
```
(On NVIDIA hosts, prefix with `CUDA_VISIBLE_DEVICES=0`.)

This covers fresh scoring, serial prefix-cache reuse, and parallel shared-state scoring. Reproduce the native reranker measurements separately:

```bash
CUDA_VISIBLE_DEVICES=0 python benchmarks/shape777_reranker.py \
python benchmarks/shape777_reranker.py \
--model Qwen/Qwen3-Reranker-4B \
--revision 22e683669bc0f0bd69640a1354a6d0aebcfeede5 \
--input benchmarks/data/shape777.jsonl \
Expand Down Expand Up @@ -115,15 +118,17 @@ Regenerate the row-level predictions with the published scorer paths. The commit
score_set () {
input=$1
stem=$2
CUDA_VISIBLE_DEVICES=0 semif-score --mode serial \
semif-score --mode serial \
--model Qwen/Qwen3.5-4B \
--revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
--input "$input" --output "direct-$stem.jsonl"
CUDA_VISIBLE_DEVICES=0 semif-score --mode reranker \
semif-score --mode reranker \
--model Qwen/Qwen3-Reranker-4B \
--revision 22e683669bc0f0bd69640a1354a6d0aebcfeede5 \
--input "$input" --output "reranker-$stem.jsonl"
}
# On NVIDIA hosts, prefix the semif-score lines with CUDA_VISIBLE_DEVICES=0.
# Use --device mps|cpu and --dtype float16|float32 to override the automatic choice.

score_set benchmarks/data/authored144.jsonl authored144
score_set "$OUT/wanli256.jsonl" wanli256
Expand Down
52 changes: 42 additions & 10 deletions benchmarks/decision_vs_generation.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,13 @@
import time
from pathlib import Path

from semif_phase1.core import load_causal_model
from semif_phase1.core import (
describe_hardware,
load_causal_model,
peak_memory_bytes,
reset_peak_memory_stats,
synchronize,
)
from semif_phase1.shared import score_shared


Expand Down Expand Up @@ -78,8 +84,7 @@ def run_generation(model, tokenizer, state: str, rows: list[dict], max_new_token
streamer=streamer,
use_cache=True,
)
if next(model.parameters()).device.type == "cuda":
torch.cuda.synchronize()
synchronize(next(model.parameters()).device)
total = time.perf_counter() - started
generated = output[0, input_tokens:].detach().cpu().tolist()
text = tokenizer.decode(generated, skip_special_tokens=True)
Expand Down Expand Up @@ -115,15 +120,32 @@ def main() -> None:
parser.add_argument("--output", type=Path, required=True)
parser.add_argument("--repeats", type=int, default=3)
parser.add_argument("--max-new-tokens", type=int, default=128)
parser.add_argument(
"--device",
choices=("auto", "cuda", "mps", "cpu"),
default="auto",
help="Accelerator to use; auto selects CUDA, then Apple Metal (MPS), then CPU.",
)
parser.add_argument(
"--dtype",
choices=("auto", "bfloat16", "float16", "float32"),
default="auto",
help="Model dtype; auto prefers bfloat16 with fallback off CUDA.",
)
args = parser.parse_args()
if args.output.exists() or args.repeats < 1 or args.max_new_tokens < 1:
parser.error("Output must be new and numeric limits must be positive")
rows = [json.loads(line) for line in args.input.read_text().splitlines() if line.strip()][:21]
if len(rows) != 21 or len({row["state"] for row in rows}) != 1:
parser.error("Input must begin with one complete 21-question shared-state group")

model, tokenizer, metadata = load_causal_model(args.model, args.revision)
import torch
model, tokenizer, metadata = load_causal_model(
args.model,
args.revision,
device=None if args.device == "auto" else args.device,
dtype=None if args.dtype == "auto" else args.dtype,
)
device = next(model.parameters()).device

# Warm both paths; warmup is excluded from every reported duration.
score_shared(model, tokenizer, rows, metadata)
Expand All @@ -134,15 +156,24 @@ def main() -> None:
direct_runs = []
direct_outputs = None
for _ in range(args.repeats):
torch.cuda.reset_peak_memory_stats()
reset_peak_memory_stats(device)
direct_outputs, timing = score_shared(model, tokenizer, rows, metadata)
direct_runs.append({**timing, "peak_cuda_bytes": torch.cuda.max_memory_allocated()})
peak = peak_memory_bytes(device)
direct_runs.append(
{
**timing,
"peak_cuda_bytes": peak if device.type == "cuda" else None,
"peak_memory_bytes": peak,
}
)

generation_runs = []
for _ in range(args.repeats):
torch.cuda.reset_peak_memory_stats()
reset_peak_memory_stats(device)
run = run_generation(model, tokenizer, rows[0]["state"], rows, args.max_new_tokens)
run["peak_cuda_bytes"] = torch.cuda.max_memory_allocated()
peak = peak_memory_bytes(device)
run["peak_cuda_bytes"] = peak if device.type == "cuda" else None
run["peak_memory_bytes"] = peak
generation_runs.append(run)

direct_choices = [
Expand All @@ -156,7 +187,8 @@ def main() -> None:
report = {
"version": "decision-vs-compact-generation-v2",
"model": metadata,
"hardware": torch.cuda.get_device_name(0),
"hardware": describe_hardware(),
"device": device.type,
"input": {
"path": str(args.input),
"sha256": hashlib.sha256(args.input.read_bytes()).hexdigest(),
Expand Down
37 changes: 31 additions & 6 deletions benchmarks/shape777.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,12 @@
import time
from pathlib import Path

from semif_phase1.core import load_causal_model
from semif_phase1.core import (
describe_hardware,
load_causal_model,
peak_memory_bytes,
reset_peak_memory_stats,
)
from semif_phase1.direct import score
from semif_phase1.serial import SerialPrefixScorer
from semif_phase1.shared import score_shared
Expand All @@ -23,6 +28,18 @@ def main() -> None:
parser.add_argument("--input", type=Path, required=True)
parser.add_argument("--output", type=Path, required=True)
parser.add_argument("--max-tokens", type=int, default=4096)
parser.add_argument(
"--device",
choices=("auto", "cuda", "mps", "cpu"),
default="auto",
help="Accelerator to use; auto selects CUDA, then Apple Metal (MPS), then CPU.",
)
parser.add_argument(
"--dtype",
choices=("auto", "bfloat16", "float16", "float32"),
default="auto",
help="Model dtype; auto prefers bfloat16 with fallback off CUDA.",
)
args = parser.parse_args()
if args.output.exists():
parser.error("Output must be new")
Expand All @@ -32,26 +49,32 @@ def main() -> None:
groups[row["group_id"]].append(row)
if len(rows) != 777 or len(groups) != 37 or any(len(group) != 21 for group in groups.values()):
parser.error("Expected the committed 37-state x 21-question fixture")
model, tokenizer, metadata = load_causal_model(args.model, args.revision)
import torch
model, tokenizer, metadata = load_causal_model(
args.model,
args.revision,
device=None if args.device == "auto" else args.device,
dtype=None if args.dtype == "auto" else args.dtype,
)

first = next(iter(groups.values()))
score(model, tokenizer, first[0], metadata, args.max_tokens)
warm_serial = SerialPrefixScorer(model, tokenizer, metadata, args.max_tokens)
for row in first:
warm_serial.score(row)
score_shared(model, tokenizer, first, metadata, args.max_tokens)
device = next(model.parameters()).device
report = {
"version": "shape777-published-v1",
"input_sha256": hashlib.sha256(args.input.read_bytes()).hexdigest(),
"model": metadata,
"hardware": torch.cuda.get_device_name(0),
"hardware": describe_hardware(),
"device": device.type,
"timing_scope": "Warm model; includes prompt construction, tokenization, transfers, forward passes and CPU readout.",
"results": [],
}
predictions = {}
for mode in ("fresh", "serial_prefix", "parallel_shared"):
torch.cuda.reset_peak_memory_stats()
reset_peak_memory_stats(device)
started = time.perf_counter()
values, state_times = [], []
for group in groups.values():
Expand All @@ -67,13 +90,15 @@ def main() -> None:
state_times.append(time.perf_counter() - mark)
elapsed = time.perf_counter() - started
predictions[mode] = values
peak = peak_memory_bytes(device)
report["results"].append(
{
"mode": mode,
"wall_seconds": elapsed,
"decisions_per_second": len(values) / elapsed,
"state_p50_seconds": statistics.median(state_times),
"peak_cuda_bytes": torch.cuda.max_memory_allocated(),
"peak_cuda_bytes": peak if device.type == "cuda" else None,
"peak_memory_bytes": peak,
}
)
reference = {row["id"]: row for row in predictions["fresh"]}
Expand Down
38 changes: 32 additions & 6 deletions benchmarks/shape777_reranker.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,13 @@
import time
from pathlib import Path

from semif_phase1.core import load_causal_model, softmax
from semif_phase1.core import (
describe_hardware,
load_causal_model,
peak_memory_bytes,
reset_peak_memory_stats,
softmax,
)
from semif_phase1.reranker import score_pair_batch


Expand All @@ -31,6 +37,18 @@ def main() -> None:
parser.add_argument("--output", type=Path, required=True)
parser.add_argument("--pair-batch-sizes", default="1,4,8")
parser.add_argument("--max-tokens", type=int, default=4096)
parser.add_argument(
"--device",
choices=("auto", "cuda", "mps", "cpu"),
default="auto",
help="Accelerator to use; auto selects CUDA, then Apple Metal (MPS), then CPU.",
)
parser.add_argument(
"--dtype",
choices=("auto", "bfloat16", "float16", "float32"),
default="auto",
help="Model dtype; auto prefers bfloat16 with fallback off CUDA.",
)
args = parser.parse_args()
if args.output.exists():
parser.error("Output must be new")
Expand All @@ -43,16 +61,22 @@ def main() -> None:
groups[row["group_id"]].append(row)
if len(rows) != 777 or len(groups) != 37 or any(len(group) != 21 for group in groups.values()):
parser.error("Expected the committed 37-state x 21-question fixture")
model, tokenizer, metadata = load_causal_model(args.model, args.revision)
import torch
model, tokenizer, metadata = load_causal_model(
args.model,
args.revision,
device=None if args.device == "auto" else args.device,
dtype=None if args.dtype == "auto" else args.dtype,
)

warm = next(iter(groups.values()))[0]
score_pair_batch(model, tokenizer, [(warm, option) for option in warm["options"]], args.max_tokens)
device = next(model.parameters()).device
report = {
"version": "shape777-reranker-published-v1",
"input_sha256": hashlib.sha256(args.input.read_bytes()).hexdigest(),
"model": metadata,
"hardware": torch.cuda.get_device_name(0),
"hardware": describe_hardware(),
"device": device.type,
"semantic_contract": (
"Two independent yes/no relevance passes per binary decision; "
"option log-odds normalized only for relative comparison."
Expand All @@ -61,7 +85,7 @@ def main() -> None:
}
prediction_lines = []
for size in sizes:
torch.cuda.reset_peak_memory_stats()
reset_peak_memory_stats(device)
started = time.perf_counter()
state_times, predictions = [], []
forward_seconds = padded_tokens = 0
Expand Down Expand Up @@ -91,6 +115,7 @@ def main() -> None:
)
state_times.append(time.perf_counter() - state_started)
elapsed = time.perf_counter() - started
peak = peak_memory_bytes(device)
record = {
"pair_batch_size": size,
"judgments": len(predictions),
Expand All @@ -100,7 +125,8 @@ def main() -> None:
"state_latency_p50_seconds": statistics.median(state_times),
"state_latency_p95_seconds": percentile(state_times, 0.95),
"padded_tokens": padded_tokens,
"peak_cuda_bytes": torch.cuda.max_memory_allocated(),
"peak_cuda_bytes": peak if device.type == "cuda" else None,
"peak_memory_bytes": peak,
}
report["results"].append(record)
prediction_lines.extend({"pair_batch_size": size, **row} for row in predictions)
Expand Down
2 changes: 1 addition & 1 deletion docs/METHOD.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,7 +39,7 @@ Qwen3-0.6B, MiniCPM5-2B, and Qwen3.5-4B use the same frozen prompt and native BF

An owned fixture contains 37 states and 21 fixed binary criteria per state, giving 777 decisions. States are roughly 8,000 characters and exercise repeated-context computation. It matches the count geometry of the public Every/Jev demonstration, but does not reproduce its unpublished documents, token lengths, hardware, API path, or model. Therefore it is a systems measurement, not a Jev head-to-head benchmark.

Direct modes are fresh batch-one scoring, serial suffixes after one state prefill, and parallel suffix branches after one state prefill. The reranker repeats the state for two independent yes/no option pairs per binary decision and tests ordinary pair batching. Timings use one RTX 3090 with a warm-loaded BF16 model and include prompt construction, tokenization, transfers, forward passes, and CPU readout; model loading and result-file writes are outside the timed region.
Direct modes are fresh batch-one scoring, serial suffixes after one state prefill, and parallel suffix branches after one state prefill. The reranker repeats the state for two independent yes/no option pairs per binary decision and tests ordinary pair batching. Published timings use one RTX 3090 with a warm-loaded BF16 model and include prompt construction, tokenization, transfers, forward passes, and CPU readout; model loading and result-file writes are outside the timed region. The scorer also runs on Apple Metal (MPS) and CPU via automatic device selection (CUDA, then MPS, then CPU) with BF16 preferred and float16/float32 fallback off CUDA; the fallback dtype is recorded in `model.dtype`. MPS/CPU runs are functionally equivalent but their timings and peak-memory fields are not comparable to the published CUDA figures.

## Interpretation rules

Expand Down
19 changes: 18 additions & 1 deletion src/semif_phase1/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,18 @@ def main() -> None:
parser.add_argument("--input", type=Path, required=True)
parser.add_argument("--output", type=Path, required=True)
parser.add_argument("--max-tokens", type=int, default=4096)
parser.add_argument(
"--device",
choices=("auto", "cuda", "mps", "cpu"),
default="auto",
help="Accelerator to use; auto selects CUDA, then Apple Metal (MPS), then CPU.",
)
parser.add_argument(
"--dtype",
choices=("auto", "bfloat16", "float16", "float32"),
default="auto",
help="Model dtype; auto prefers bfloat16 with float16/float32 fallback off CUDA.",
)
args = parser.parse_args()
if args.output.exists() or args.max_tokens < 1:
parser.error("Output must be new and max-tokens must be positive")
Expand All @@ -29,7 +41,12 @@ def main() -> None:
parser.error("Input is empty")
for row in rows:
validate_row(row)
model, tokenizer, metadata = load_causal_model(args.model, args.revision)
model, tokenizer, metadata = load_causal_model(
args.model,
args.revision,
device=None if args.device == "auto" else args.device,
dtype=None if args.dtype == "auto" else args.dtype,
)
args.output.parent.mkdir(parents=True, exist_ok=True)
with args.output.open("x") as destination:
if args.mode == "shared":
Expand Down
Loading